Feature request: would it be possible to have customizable Fressian handlers for pss on-disk storage? Right now they seem to be a https://github.com/replikativ/datahike/blob/main/src/datahike/index/persistent_set.cljc#L554-L564.
Case in point – some of the values in my db are cljc.java-time time values (instances of LocalDateTime etc). For now my workaround is to replicate Datahike’s (defmethod di/add-konserve-handlers :datahike.index/persistent-set) and add my handlers there, but this is somewhat unwieldy...
If the set of serializable value types is closed by design (interoperability), that’s understandable, but ideally I’d like to see transactions fail early when an unserializable value is attempted to be transacted. Right now the error is somewhat opaque.
@dj942 I see, I was looking into opening the value types, but it is very easy to use the database in the wrong way this way, and yes it harms replication, too. So for now I will keep it closed to not have responsibility for that, but I added the store-ref type to store blobs. I think I should revisit cljc.java-time and/or Java Instant time support as well, so far I kept it simple and portable though. If you could give me a minimum reproducible example for the error and how you would like to use it that would be helpful.
I think a better way might be to add support for custom types to https://github.com/replikativ/datahike/blob/main/doc/unstructured.md. That way you make sure your custom types are properly translated into datahike's index layout and you really benefit from it (unless an unindexed blob is all you need). Ultimately this should become an option to add in the standard (transact conn {:tx-data ... :unstructured ...}) I think. Or a custom map wrapper for :tx-data itself.
For types that are compare adding Fressian handlers and new db types might be useful in general, since they would be indexed as is.,I am just worried about how to maintain Datahike for an open-ended system like this.
Thanks, this looks useful!
For a type like Instant though (which has internal structure (a pair of long and int) but is conceptually an atomic value), I guess I’d rather serialize it to a string before transacting than use the implicit schema inference. It’s doable, it’s just some additional friction.
I guess what I’d really welcome is native support of date/time types in Datahike, not necessarily arbitrary extensibility.
Understood, we should figure this out. I hope I can get to it soon, if not, please keep annoying me 🙂
Wrote a blog post that might be food for thought on this: https://blog.danieljanus.pl/datomic-map-values/
I noticed a weird behavior when trying to do gc-storage for a db that has no history enabled. This works: @(d/gc-storage (:datahike ctx) (java.util.Date.)) But this does not work: @(d/gc-storage (:datahike ctx)) The second version does not release any blocks. Is this intended behaviro?
The date is a cutoff date, by default it is the beginning of time, so then the gc will only collect dead branches (which you don't seem to have). I would use a cutoff of at least an hour or day or a week, depending on how long your queries run and you want to access old snapshots (you have them in the history db as well, too, though). It is mostly about concurrent distributed readers that might still access these databases and if you want to use the git-like branching features.
My datahike db was 7 GIG and by running gc-storage without a date, it left it at 7 GIG. Then I ran gc-storage with todays date, and it compacted it to 60 MB. The db does not use history. I use a single machine with a single clojure instance, which gets auto-restarted every Sunday night. So I am really surprised that it did not clean anything without specifying a date.
I dont understand the access to old snapshots when I have history disabled. Can you please elaborate.
@alekcz360 and I discussed adding automatic background gc a year ago, this is here now https://github.com/replikativ/datahike/blob/main/doc/gc.md#background-gc-mode, but not yet automatically started. If you can test it and lmk whether it works well that would be cool. I want to make this a stable solution that is turned on/off by a config setting automatically.
@whilo can you explain what are snapshots? When does a snapshot happen? Are old snapshots snapshots that are not the most recent one? In terms of @(d/gc-storage (:datahike ctx) (java.util.Date.)). So if I do this and there is a query that is running from before calling gc-storage, then the query would return wrong data?
I have implemented a cron job (that runs inside my app, so a pure clj based function), that will run every sunday 3am. I will report if there are errors after I implemented this feature. Does this help? Or you want only automatic gc report?
Datahike has two forms of history, internal history indices, and a commit-like history (which is equivalent to as-of db). On each write it writes a commit and updates the branch head, the commit is an immutable snapshot. In a setting without diff-buf it contains at least one path from root to the branch in each index tree https://hypirion.com/musings/understanding-persistent-vector-pt-1. With the new settings it is in most cases just a single write without the full path, reducing the snapshot size. These snapshots are what readers operate on to be free from coordination.
I would advice against a cron job, you should try to keep all write operations on Datahike in a single machine, I assume you connect to your normal writer and dispatch the gc in the cron job, which is fine, but if you would spawn your own parallel writer then this could lead to data loss.
Just to be clear, because this is a sensible topic. I run periodically @(d/gc-storage datahike-conn (java.util.Date.)). datahike-conn is also used for queryies and transactions. All on one machine in one process. Is this safe?
Yes
If you use now as the cutoff date, which you do here, make sure to use one of the latest versions though, because in this case there was actually a race with the transactor which made it unsafe in some circumstances.
I would in general advice to not use (java.util.Date.) unless you don't want to run concurrent queries.
Because the gc will immediately invalidate all readers older than (java.util.Date.), it might just take it long enough to get there.
Use something like (- (java.util.Date.) (* 1 60 60 1000)) [one hour].
Thanks @whilo the GC at least for me is a huge improvement. I so far had to once a month dump a local filesystem db and then import to a new db to keep the growth of the db in check. With periodic GC it becomes something that can run on unattended servers for a long time. Huge plus!!!