To make Datahike fully plannable for production we need to provide the ability to bound string and byte array input sizes. To do so I want to merge this PR in the next days https://github.com/replikativ/datahike/pull/861 , feedback is welcome.
I opted for Datomic defaults (mostly for compatibility of 4096 char length max), but it is opt-out and existing databases will not be affected, only newly created ones with schema-on-write. I have heard people complain about fixed character sizes, but there is a reason why Postgres (VARCHAR) and Datomic etc. bound sizes by default, otherwise you can blow through all cache estimates and storage segments sizes without realizing. @pedrogbranquinho also contributed https://github.com/replikativ/datahike/pull/859 (slightly modified by me), which already bounds our query cache size properly, in combination with input size bounds this then should provide total bounds for expected memory usage. There might be some other spots where we still need to limit things though.
If I remember correctly, we stored EDN documents that are bigger than 4096 chars into Datomic (https://docs.datomic.com/schema/schema-reference.html#notes-on-value-types). It caused all kind of problems like CPU-spikes (due to Fressian decompression, I guess) and premature cache evictions. So in general I agree, it should be limited (but it will make migrations from Datomic Pro to Datahike more difficult). However, I think there should be an option in the schema to define that the value should be stored separately in the underlying konserve storage. Or maybe it could be a different valueType. Then the datom v could be an UUID that can be used in the konserve storage to lookup the large value.
@maxweber This is very good to know. I did not know that Datomic Pro was not enforcing the limit. I could also keep it off by default, but I guess being able to turn it off without much hassle should be good enough.
@whilo I would keep it off by default, otherwise it is breakage for all existing Datahike application or?
Existing databases would keep it off, newly created ones would have it on. Maybe you are right, but then I can never turn it on by default.
But even if only newly created ones would have it on by default, it is light breakage, since people may re-import data into new dbs when doing migrations
Yes, I could also warn for now that it might be turned on in the future by default and that people should turn it on.
Policy: Never break existing users unless there is no other way to fix something critical. @maxweber This is merged now and prints a warning when you create a db without limits.
Not ideal for new users though, but I think they will pick it up. I still need to update the docs everywhere to include it.