Datahike 0.8.1705 is released, turning the https://github.com/replikativ/datahike/blob/main/doc/query-engine.md on by default.
Over the last two months I have been using it in cljs (where is already default, but not fully mirroring Clojure yet), pg-datahike and several projects. I applied fixes to it on the way, but overall it has proven very valuable, in particular for pg-datahike.
You can turn it off by setting *disable-planner* to true or by setting DATAHIKE_QUERY_PLANNER=false as an environment variable.
Please lmk and/or open issues if you run into any!
https://github.com/replikativ/datahike/pull/836 I want to land to reduce write amplification and reintroduce the hitchhiker-tree ideas, after which I would feature freeze Datahike for its 1.0 release. Then I will prioritize bug fixes, refactorings to unify Clojure/cljs and turn our secondary indices and maybe some other experimental features stable.
For bug fixes, do you have a list in mind? Was going through the issues list in the datahike repo earlier, so I could help out if it's alright.
I just also realize that the dthk tool does not include the migration, because we have not yet put the migration import/export into the datahike.api. @alekcz360 has quite a bit of experience for import/export and i think also continuous replication (we also support this now transparently with konserve-sync). I think we might want to think this through a little before exposing it in the stable API or exposing more machinery. I am in general somewhat conservative when it comes to adding features that need to be maintained in the long run, but this makes sense to have as part of the stable API (which at the moment is one flat specification). We also picked CBOR there to make it easy to load, process and emit Datahike databases from other languages, but maybe this is not the optimal way.
Thanks for all of the pointers. I wanted to get more comfortable with native clojure, so the reflection and build issues sounds interesting.
Cool. Lmk if you have questions or need anything.
if you want to test a release you have to change the config and push to your own fork and add the secrets for your personal clojars account
@whilo Whenever you have free time, native build fixes PR above is ready for review.
bb test bb-pod is a remaining failure in the ni workflow, but I need the latest CI/CD logs to start tackling that
@hello571 I https://github.com/replikativ/datahike/pull/851#issuecomment-4885039087 in the PR, yes the bb-pod is the remaining error. Unfortunately I had to manually mirror the PR internally, which I think can trigger a native image release and this is kind of the only way to do this, right @timok? @hello571 Can you not reproduce the bb-pod error locally without the CI/CD?
I probably miss something 🙂
> Can you not reproduce the bb-pod error locally without the CI/CD? @whilo Actually, I already did. But, since it was in core logic-space than just build stuff wanted to double check first.
FAIL in (pod-workflow) (/Users/harismh/Developer/datahike/./bb/resources/native-image-tests/run-bb-pod-tests.clj:22)
as-of timestamp
expected: (= #{[2 "Bob" 30] [1 "Alice" 20] [3 "Charlie" 40]} (d/q (quote [:find ?e ?n ?a :where [?e :name ?n] [?e :age ?a]]) (d/as-of (d/db conn) timestamp)))
actual: (not (= #{[2 "Bob" 30] [1 "Alice" 20] [3 "Charlie" 40]} #{[2 "Bob" 30] [1 "Alice" 20]}))Looks like the CI error is the same too. I'll fix this in the same PR if that works.
Perfect, thanks.
Merged. Thanks @hello571! Lmk if there are other things you are interested in and we can also work on those together.
@whilo Thanks for the assist. Open to any suggestions. Since 1.0 is still the focus, I could take a look at more bug reports/edge cases if you know any. By the way, there are still some reflection warnings that aren't halting builds: https://github.com/replikativ/datahike/actions/runs/28857461570/job/85594966757#step:6:745
Fixing them is definitely a good idea. You can also lmk what you are personally interested in and we can discuss this, too.
It would be fun also to look into some of the old papers from https://en.wikipedia.org/wiki/Fifth_Generation_Computer_Systems again. Maybe some ideas can be made work now much better together with LLMs and modern techniques 😉
core.logic for instance has a Datomic integration, which should be trivial to port for Claude and allows it to search over Datahike databases effectively.
Obviously this is a bit speculative and maybe not directly application related, but I do think there are new possibilities in this space.
Ideally one could use for instance the dvergr agent harness, let it systematically download the papers if it can get them (or maybe they can be downloaded in bulk), then index them with different extraction methods as well through the LLMs and make them themselves searchable in a Datahike db with embedding vectors through proximum and meta data / full-text matches for instance. That way one could then have coding assistants in the REPL use this logic programming database (and maybe others), to systematically study ideas from the literature and implement them in the same system. This is a bit of a stretch, but I think a lot less crazy than it sounds.
On the practical side we need to make sure Datahike works really well in the browser. For the 1.0 release ideally I want to graduate cljs support to be production ready. I need to port over most tests to deftest-async, but it also just needs to be tested quite a bit, including performance and reliably. Also probably the kabel-writer network sync robustness under connection loss etc.
Testing secondary indices in general is also a good idea, they also should graduate to being production ready for 1.0, I think.
> Fifth Generation Computer Never knew about this, but it's fascinating. Thanks for the share. I only read about the Memex beforehand, but this logic-first approach is fundamentally different. Same fate as the Lisp Machine though... > I need to port over most tests to deftest-async, If this is for nodejs_test.cljs, could take a look at this. Wonder if for t/run-tests we should try adding all .cljc ns and make patches along the way, eventually dropping that call and running the full suite on both runtimes. As currently, it looks like this call is unused:
(defn ^:export test-cljs []
(datahike.test.core/wrap-res #(t/run-all-tests #"datahike\..*")))
> core.logic for instance has a Datomic integration, which should be trivial to port for Claude
New to core.logic, but definitely sounds interesting. Would be outside my comfort zone if you're willing to entertain my naive PRs haha. Initially thinking of a new ns under experimental. For tests, I imagine it would be proving queries like this https://github.com/clojure/core.logic/blob/master/src/main/clojure/clojure/core/logic/datomic.clj#L135-L140I'll make a PR for the remaining reflection warnings soon.
@whilo Making progress on the native builds fixes. Was curious about some of the requiring-resolve calls. Are they for lazy loading, getting past some circular requires, both? Dropping some of those requiring-resolve calls in favor of reader conditionals does seem to fix the native builds. But, was wondering if that was the right approach.
Nice. In general they should be removed, I must have missed them. Can you point me to them?
Claude injects them automatically and I usually do an additional pass to remove them, I will have to manage this better.
This is my bad, I can also do it, you don't have to cleanup this Claude cruft for me.
No worries, they were in query.cljc. Current draft PR: https://github.com/replikativ/datahike/pull/851
If you could run CI/CD on that branch, that would be helpful to see the latest build failures