Is there an intended way to evolve a schema within datalevin on a production database?
Specifically; if I define a schema and a way to do migrations, is the intended way to open the connection with a "v1" schema and then call update-schema on it repeatedly? Or should I open the connection with the current schema as it stands? If I'm supposed to open it with the current schema, how do I detect that? Do I have to store migration state out of band to determine what schema to use?
Alternatively, am I supposed to just open the schema with the most recent schema and datalevin will just handle any new attributes added as long as I don't rename attributes?
The book gave suggestions on the functions used for updating a schema, but little in the way of guidance.
It's really depending on what you mean by "migration".
where you have a series of "migrations", traditionally sql scripts or code that you run to set up a sql database to have the correct shape. and htne you can apply the migrations necessary that haven't been run yet, in the order specified, to get the database to have any expected changes
it's a means of applying changes over time to a sql database's schema in a way that's repeatable (you need to set up a new version of the db) and atomic (you and someone else are making changes to the database schema at the same time, how do you make sure you get each other's changes too?) and idempotent (you generally keep track of which migrations have been run in their own table (through a library or manually) so you only run the migrations that haven't been applied yet)
i know that datalevin is "schemaless", but still seems worthwhile to have some sort of guidance on handling asynchronous changes over time to the shared database
I want to be clear; I am very familiar with migrations, and I have written a ton of them and even written migration frameworks in my time.
What I am specifically curious about is how datalevin handles mismatches between declared schema and the actual stored database, both in cases where the declared schema contains fewer attributes than the db, and in cases where the schema contains more attributes than the db, as well as to understand the expectations around when applying migrations what the actual state transitions of the db connection need to be wr/t the use of update-schema compared to the static schema passed in during construction of the connection.
In particular:
• is update-schema idempotent?
• is it safe to initialize a connection to an existing datalevin store with a schema different than the one it was created with?
◦ Is it safe to connect to a datalevin store with declared attributes which are not yet present in the store?
◦ Is it safe to connect to a datalevin store with a schema that only describes the migration storage data, and then use update-schema to apply schema migrations not yet applied to the db and to sync the current attribute set with the attributes that exist currently?
▪︎ What does update-schema even do when you ask to add an attribute that wasn't in your declared schema but is already present in the store?
update-schema can do a lot. Attribute addition/change is idempotent. However, attribute deletion and rename, by definition, cannot be.
In general, Datalevin does things the way one would expect. The goal is to have good ergonomics, that includes minimizing surprises.
The stored schema and the opening schema are merged, as you would expect in a Clojure map operation. So if you already have data in DB, and you open it with a completely incompatible schema, you will get errors, as expected.
Datalevin currently only does safe migration automatically, e.g. untyped -> typed, but not reverse.
Is it safe to connect to a datalevin store with declared attributes which are not yet present in the store? Yes. Of course. You can have an empty DB with a schema. No problem.
Is it safe to connect to a datalevin store with a schema that only describes the migration storage data, and then use update-schema to apply schema migrations not yet applied to the db and to sync the current attribute set with the attributes that exist currently? Yes, of course. There is stored schema that reflects the current data. So you can open a db with an empty schema, it will work just fine. We do schema merges on open.
In Datalevin, schema is meta data, not data. Datalevin is more similar to SQL db than Datomic, in this regard.
Awesome, thanks for the detailed answer! That is extremely helpful.
I have made update-schema fully idempotent, available in master branch.