Just completed my first full import of the English DBPedia with 13 million entities and 224 million facts in 10 hours (6 gig heap + off heap LMDB usage, 27 gib final database size). This is with a few improvements to help bulk throughput and with an online gc to avoid accumulating unnecessary snapshots (not yet released). I am thinking about providing this as a download through IPFS, ideally through yggdrasil, if people are interested (including source code to reproduce).
Interesting, I may use yggdrasil to model L1 (low-level abstraction memories) in the Knowledge Graph (KG) module, in hive-mcp.
It will be useful because LLMs won't need to re-read file-system constantly. They will just know where things are, with certain degree of certainty. And thus provide token optimization. Of course, will eventually reread certain files when degree of certainty decreases below threshold
It would be compatible with datalevin, right? Certain nodes link to the memory-system in yggdrasil
Hey @pedrogbranquinho. Nice work on hive-mcp! I have been working on a PKM backbone shared between people and LLMs as well a bit (see also #simmis (WIP)). Datalevin is a mutable database and does not retain snapshots or allow branching and merging. It is very well engineered for performance and memory efficiency on top of LMDB, but because of the lack of this persistent memory model it does not fit into yggdrasil, does not match Clojure's persistent memory semantics, and does not enable the wider program/simulation synthesis/inference frame that I am working towards. This is why I am pushing Datahike atm. (see for instance the DBPedia import above, which is a great backbone for other knowledge bases).
I see! I will definitely take a look at datahike. Everything was written with SOLID principles, so it's easy to change gears when needed. Currently, I will use yggdrasil for grounding knowledge of data state etc. Which can be done apart of the rest of KG that currently is stored by Datalevin layer. But, good to get this comparison. Thanks a lot! If you want to colab on anything, or maybe you see how hive-mcp would/could fit into simmis... There is a module called Olympus that currently only have the backend done in clojure (it's an engine to arrange lings [workers] into a screen partitioning) and the UI could be done in multiple plug-gable places (web, Emacs, etc). Seems to be a nice initiative 👏