I have a much larger idea here, but I think it might be productive to break it into parts. So part 1/N:
A repository today serves two roles.
One is to index artifacts for consumption. I tell a repository that I have a variant of dev.mccue/json and that repository has some policy by which they decide whether or not to list that variant for others to consume. Usually this is "I prove that I own the domain name and that gives me unique permission to publish artifacts with that group id." So the repository decides what the authoritative dev.mccue/json v1.2.3 is.
The other is to serve artifacts as a CDN. I ask clojars for dev.mccue/json v1.2.3 and it will serve me the bytes for that artifact. This is the part of a repository that takes the most resources, bandwidth and storage are not free.
So the first part of the idea is to decomplect these two responsibilities. As a strawman (and I am trying to nerd snipe you @tcrawley but anyone else please chime in), what if when I asked a repository for dev.mccue/json v1.2.3 instead of getting the actual bytes of the artifact I got the SHA-256 of that artifact? This reflects the "indexer"'s opinion of what that library is. Then, as a second step, I use that SHA-256 to get the actual bytes of the artifact.
Importantly, this would mean that the indexer does not necessarily need to be the CDN. You should be able to use the SHA-256 to find the artifact from any number of sources. (IPFS, Bittorrent, an internal blob store, whatever).
Thoughts?
What is the problem that this solves? The cost of running a blob store? A key promise Clojars makes is to keep artifacts around forever so old builds can keep working. If we outsource that to IPFS or otherwise then we lose that guarantee (I think?)
It's an interesting idea, but I agree that we should talk about problems first. And I'll note that it would mean that clojars would no longer be a maven repository, meaning no existing tooling could use it.
I think it's hard to strike a balance between talking about an ideal system and figuring out what we could do in principle. And I struggle with that which is part of why I didn't talk too much about how this would affect clojars specifically
I think there are some clever ways you can do everything I'm thinking and still have clojars be a maven repository
(stuff like just pass a header that says where your blobs are and clojars redirects there instead of to its own buckets)
And there are several problems I am thinking about. The biggest one that is relevant for this specific design is that to be a repository you both need to perform the role of a curator and you need to be a globally available CDN for everybody's builds, and you need to be an internet archive style permanent store of artifacts
That combination of responsibilities shapes what repositories exist and can be relied on. People rely on Maven Central because it does domain verification, has been up for long enough the people trust it to continue to be up, and has historically served as a globally available CDN for their builds
It also means I can't host my own repository without taking on all of those responsibilities
Okay maybe this is a good way to put it - clojars right now serves the role of • "indexer/curator" - it decides what libraries is it lists and who is allowed to provide them • "CDN" - it provides these libraries to builds any company is running for free (while begging them to not make use of this part because that cost a lot please set up a mirror and most of the time people don't) • "Archiver" - if someone published a library it will be on clojars forever. So you could push it on there and be confident you won't lose it so long as you're confident in the health of the archive
And I think every repository really just wants to play the first and the last role. Nobody wants to be providing unlimited downloads to every startup in the country, but it's weirdly impractical to tie operating income to traffic in that way
Of course you can draw a circle around all of it and be an one org that does it all, but (see my first point about I have trouble detaching designing a system as I would want it ideally from talking about systems as they exist today) that sucks
Also because we tie the roles of curator and archiver, our repositories have "the Facebook problem." You don't use Facebook because you like Facebook you use Facebook because everyone else uses Facebook. You put your libraries where everybody else puts their libraries because doing otherwise would be impractical
But if you draw that line out enough you end up with a single repository, meaning a single organization providing free CDN services to the entire planet. The fact that most people use clojars with Maven Central brings us to two but it's still a rounding error
The exercise for me is trying to pick things apart enough that we can put them back together in a different shape. I'm trying to come up with "what we would do if we were starting from a blank slate" and then working backwards once we get there.
My goal with separating indexers/curators from cdns/blob stores/archives is in part just to say "yes we could do this in principle" because having those things be separate is a prerequisite for the next slide, which is making it so that repositories don't need to be the source of Truth for the identity of their publishers
And this might sound a lot closer to what I was ranting about at the after party. Right now it is the responsibility of every repository to do their own group id allotments. Different repositories will do this under different policies (which does suck from a security point of view because you need to trust every single mirror and alternate repository you add to your build. You need manual intervention to make it so that domain hijacking isn't an issue, handling when people abandon their projects/actually die is rough, etc.)
And I need to start my work day and I don't know how to tie this up in a neat bow, maybe I just need to sit down and write out the whole manifesto including a "how do we cope with the current state of the world"