geschichte — Git, as a queryable database A fresh take on the original geschichte (2013, which grew into replikativ): a real Git implementation — real objects, packs, refs, and the wire protocol, interoperable with native git — but the repository is stored as datoms in Datahike. So the whole repo is one Datalog query surface: ask for every commit, path and ref, run graph algorithms over the commit DAG, and join a repo against another repo or your app's own data — in a single query.
clojure
;; authors of the merge commits in HEAD's ancestry — reachability × structure, one query
[:find ?author (count ?c)
:where [(graph/transitive-closure ?g $ ?head) [?c ...]]
[?c :geschichte.commit/parents ?p1] [?c :geschichte.commit/parents ?p2]
[(not= ?p1 ?p2)] [?c :geschichte.commit/author ?author]]
Runs on the JVM and in the browser (CLJS), ships a native ges CLI, Apache-2.0. Feedback very welcome → https://github.com/replikativ/geschichteIt uses the new blob support in datahike to store the blobs. It is going to back muschel's git implementation and filesystem backend. Git was the only native dependency we had, this makes it possible to unify dvergr's frontend compatibility, i.e. following a backend coding agent in a datahike based web frontend of simmis.
Crazy! Great work
does this have a similar intent as codeq lib ? https://github.com/Datomic/codeq/blob/master/doc/codeq.org
IIRC codeq does semantic analysis so you can query language-level stuff
Very cool! I sort of understood this as not just an analysis tool, but a means toward progressive ways of coding, where functions could be callable from a larger immutable ether. And also historic function implementations are stable and not break on new library releases. And IDE can operate on db, not just files.. more like the classic smalltalk style.
Some prior art: https://blog.phronemophobic.com/dewey-sql.html Would be cool to combine this with https://cloogle.phronemophobic.com/doc-search.html# by @smith.adriane
Imagine this rich datalog query across the entire clojure lib ecosystem think_beret
Very cool @whilo! I've been thinking about this concept for a long time (codeq and related), but never got much further then simple prototypes. One thing I thought of is normalizing forms before hashing. So for instance the following forms would be the same after normalization:
(defn foo [a] a)
(defn foo [b] b)
I believe in https://github.com/replikativ/geschichte/blob/5c3b16292e68a6ec6caf376ec812ed91c61822dc/src/geschichte/code.clj#L165-L166this normalization doesn't happen. In one of my own efforts I would walk over a form and deterministically change all local bindings. It would just rename the bindings in order of appearance so a,
b . More advanced efforts would replace symbol names with there unique id so:
(defn foo [b]
(clojure.string/replace b "foo" "bar"))
would be hashed as something like
(fn [a]
( a "foo" "bar"))
Do you think it make sense to apply this normalization step? I feel 10,090 distinct forms is still a lot for Clojure's history. I can imagine normalization would reduce it a bit more.Yes, I don't normalize yet, but I am thinking about this in general a lot (for datahike, for ansatz, for the raster compiler etc.). I am happy to explore this direction, I just wanted to represent git's semantics first.
@chromalchemy I think it would be interesting to modify the sci interpreter like this, it is also closer to the unison language, I think. It might be a good model for #C09622F337D, not sure yet.