Hello, just checking my mental model of what datahike does: if I transact this
({:product/id "foo", :details/collected-at 0, :details/content {:a 1, :b 2}}
{:product/id "bar", :details/collected-at 10, :details/content {:a 11, :b 12}})
Then this query does not return anything:
[:find ?x
:where
[?e :product/id "foo"]
[?e :details/content ?details]
[?details :a ?x]]
Is that because {:a 1, :b 2} is stored as some sort of blob?Hi @stathissideris, with Datahike you would transact each entity alone. Datahike does not work with stored documents. You should be able to return the whole document and work with it from there in your clojure code but it is not advised to do that.
ok thanks, am I correct in remembering that my example would work in datascript?
I don't know too much about datascript but I doubt it. Datahike is a fork of Datascript and I am not aware that Datascript changed that.
I don't know your schema, but normally Datascript accepts any datastructure under any attribute as long as it is not a reference or collides with multi-arity.
Datahike does accept datastructures as well and you can query it back 'as datastructure' and I believe Datascript is the same
The datastructure is put as is on the index and so you can not query inside it afaik
@filipconstantinos recently asked a similar question in the main chat. Unless you tell Datahike that the value is a ref it will not break the map down and store it as a separate entity. I think we should have a way to let Datahike transact nested JSON like this though and I am happy to discuss the design of this.
@whilo we’re on the same project as @filipconstantinos and what he was describing was different: he was reporting that although we had defined the nested map as a ref, we were getting an exception when attempting to transact as a nested map, only with the file backend.
btw I just tried the same in datascript and it doesn’t work either
Are you sure that details/content is a ref in the schema when you transact this data? I replied to @filipconstantinos in the other thread where I fixed his problem on my end this way (which is also how it should work in DataScript).
Coming back to this old thread: What we were doing (me and @filipconstantinos) was that without any schema we were transacting things like:
{:product/id i
:product/revision {:name (str "foo" i)
:prices {:per-item (rand-int 10)
:evidence :manual
:per-unit (rand-int 10)}}}
with :schema-flexibility :read
Datahike allows us to transact this, but fails under certain conditions when re-transacting similar (or even the exact same) data.
Without :schema-flexibility :read, it outright forbids the initial transaction because the nested maps have not been declared as references. I think we would be OK with having non-reference maps be stored as non-queryable blobs (maybe serialized as nippy?), but I also appreciate your solution in the main channel that involves automatic schema inference.(file backend)
@stathissideris I prototyped a JSON preprocessor for transact that can automatically infer the schema and rewrite the data for the transact DSL of Datomic/DataScript/Datahike.
(let [type->db-type (fn [v]
(cond (int? v) :db.type/long
(float? v) :db.type/float
(map? v) :db.type/ref
(inst? v) :db.type/instant
(string? v) :db.type/string
(boolean? v) :db.type/boolean))
data [{:name "Maya"
:age 62
:car {:brand "BMW"}
:likes ["bananas" "apples"]
:children [{:name "Ayia"} {:name "Tanja"}]}]
res (atom [])
temp-id (atom 0)
schema (atom [])]
((fn eval-json [data]
(cond (map? data)
(let [new-map (into {} (map (fn [[k v]]
(swap! schema conj
(cond (vector? v)
{:db/ident k
:db/cardinality :db.cardinality/many
:db/valueType (type->db-type (first v))}
(map? v)
{:db/ident k
:db/cardinality :db.cardinality/one
:db/valueType :db.type/ref}
:else
{:db/ident k
:db/cardinality :db.cardinality/one
:db/valueType (type->db-type v)}))
[k (eval-json v)])
data))
map-id (swap! temp-id dec)
new-map (assoc new-map :db/id map-id)]
(swap! res conj new-map)
map-id)
(vector? data)
(mapv eval-json data)
:else
data))
data)
{:schema (vec (distinct @schema))
:tx-data @res})
=> {:schema [#:db{:ident :name, :cardinality :db.cardinality/one, :valueType :db.type/string} #:db{:ident :age, :cardinality :db.cardinality/one, :valueType :db.type/long} #:db{:ident :car, :cardinality :db.cardinality/one, :valueType :ref} #:db{:ident :brand, :cardinality :db.cardinality/one, :valueType :db.type/string} #:db{:ident :likes, :cardinality :db.cardinality/many, :valueType :db.type/string} #:db{:ident :children, :cardinality :db.cardinality/many, :valueType :db.type/ref}],
:tx-data [{:brand "BMW", :db/id -1} {:name "Ayia", :db/id -2} {:name "Tanja", :db/id -3} {:name "Maya", :age 62, :car -1, :likes ["bananas" "apples"], :children [-2 -3], :db/id -4}]}@lidorcg If you want to pass in unstructured data a good thing to do would be to pass it first through this general json/edn interpreter, which also optionally automatically creates a schema for you. I can help you with using it, and will add official support to Datahike given your feedback. I just haven't had the time to get to it yet.
The schema enforcement of this will only work though if your data is not too heterogeneous, i.e. attributes have values that are not scalar (string, number, etc.) and collections at the same time.
This requires a bit of work before we can add it to the code base and I am not decided yet how to do it, but this is a good opportunity to solve this issue in a way that makes it easier to just throw data at Datahike without having to do a lot of upfront work even with schema on write. One issue that is missing here is that you do not want to retransact the same schema all the time.
I have an old project whose purpose was to infer clojure specs from multiple examples of the kinds of data structures in your data. The first step was to collect “stats” on each key to figure out what predicates held true for it etc. It could be used relatively easily for schema inference too: https://github.com/stathissideris/spec-provider/blob/master/src/spec_provider/stats.cljc