datahike 2026-07-30

@whilo Just wanted to say thanks for the quick turn arounds on my planner issues thanks3

2

You are welcome! I am happy to help solve issues users hit in Datahike to make sure to get it ready for 1.0.

I am also curious what you are doing with Datalog @markaddleman in general.

Happy to! The short version: Datahike is the basis for a knowledge-graph populated and queried by LLMs. The LLM extracts subject–predicate–object facts out of unstructured prose into Datahike, then reasons over them with Datalog. That’s why I keep landing on the planner around recursive rules and get-else totality specifically — I lean hard on both.

That makes sense. I have implemented a Roam-like wiki as part of #C09622F337D. Getting this ready to open source rn.

Would be happy to share work on building a good default wiki schema and functionality.

Oh very cool!

Is the new query planner+engine faster for you? I assume you used the old engine before.

Yep, I’m using the old engine right now. I’m hoping to switch to the planner once the core use case is unblocked

Thanks for the latest fix! I think the planner is completely unblocked for me now. I ran The Turn of the Screw through our extraction pipeline into a datahike graph (~12k datoms) and timed the queries:

┌────────────────────────────────────────────┬────────────────┬──────────────────┐
 │                   query                    │ ratio (OFF/ON) │     verdict      │
 ├────────────────────────────────────────────┼────────────────┼──────────────────┤
 │ coherence self-join                        │          75.7× │ planner wins big │
 │ not-join / negation                        │          2.49× │ wins             │
 │ join-heavy 5-way                           │          2.09× │ wins             │
 │ get-else provenance (statement-rows, real) │          1.47× │ wins             │
 │ aggregation (sub-2ms)                      │          0.64× │ loses (trivial)  │
 │ bounded recursive reachability (#113)      │          0.38× │ loses ~2.6×      │
 └────────────────────────────────────────────┴────────────────┴──────────────────┘
What each one is, product-side: • coherence self-join — consistency check: do the extracted claims contradict each other? (e.g. flag a character whose stated beliefs conflict) • not-join / negation — gap detection: what’s required-but-absent, or explicitly denied • join-heavy 5-way — faceted retrieval: statements matching several conditions at once (holder + attitude + polarity + topic…) • get-else provenance (statement-rows) — the workhorse: the review table listing every extracted statement with its source citation • aggregation — rollups (belief counts per character, etc.) • bounded recursive reachability — nested theory-of-mind: tracing chains of belief (“what does A believe that B believes about C”) The one loss is recursive-rule eval (the belief-chain traversal, ~2.6× slower) — wierdly, the same rule seeded in reverse wins 8–118×.

I can take a closer look at this, maybe https://github.com/replikativ/datahike/blob/main/doc/graph-algorithms.md are helpful. It would also be interesting to point Claude the recursive parts of the engine and let it experiment and compare to other recursive engines, maybe Souffle.

👍 1