Is there a good approach to represent IRIs as namespaced symbols/keywords leveraging :as-alias to keep data unambiguous and avoid having global registries? The first idea that comes to mind would be to percent-encode it. That's also what I found in discussion archives where I saw mention of as-kwi function from https://github.com/ont-app/vocabulary which implements this approach.
For example:
@base < > .
@prefix foaf: < > .
<john> a foaf:Person .
Can be turned into a keywordized triple like:
(ns example
(:require
[[http:%2F%2Fexample.org%2F :as-alias ex]
[http:%2F%2Fxmlns.com%2Ffoaf%2F0.1%2F :as-alias foaf]
[http:%2F%2Fwww.w3.org%2F1999%2F02%2F22-rdf-syntax-ns%23 :as-alias rdf]))
[::ex/john ::rdf/type ::foaf/Person]
The downside is that these namespaces are not very human-readable since slashes need to be escaped, and also the hierarchy does not match well compared to reversed domain styles like http://org.example.xyz.
Is there a better approach?I would go with the registry that you're avoiding. I have a few reasons: β’ You're not saving memory, since namespaces are just a global registry anyway, albeit one that you don't have to manage. β’ As you point out, they're more readable. β’ They can be scoped to a document. Not usually an issue since an org usually has consistent prefix mappings, but it can matter if you need to import documents. β’ The namespaces can be loaded from external resources much more easily. Yes, it's possible to create functions or macros that will do much of this for you, but loading them into the namespace mechanism would get messy. Subjectively, this "feels" like you're overloading the namespace mechanism.
Thanks, those are good practical tips.
On the conceptual level what is appealing about namespaces to me is the simplicity, it is just a mechanism of identification to globally distinguish names.
URIs complect identification, location and transport (http vs https distinction).
By document scoping you mean the # separator, right? If I understand correctly that is another inconsistency that bothers me π
When I was researching the topic in relation to URNs I found concept of https://www.w3.org/2001/tag/doc/URNsAndRegistries-50.html which looks close to the ideal in my imagination. But it does not look like there isn't any formal proposal or standard that followed.
Here is an experiment that I came up with: The above example becomes:
(ns example
(:require
[org.example.$ :as-alias ex]
[com.xmlns.$.foaf.0%2E1 :as-alias foaf]
[org.w3.www.$.1999.02.22-rdf-syntax-ns.$$ :as-alias rdf]))
[::ex/john ::rdf/type ::foaf/Person]
Approach:
- drop http/https scheme
- reverse the domain parts
- split path parts
- have placeholders to separate path and fragment (for now I went with $ for path and $$ for fragment)
- percent encode special characters (we also need to encode $ and . to avoid conflicts with separators/placeholders)
The names look pretty reasonable when using safe characters and the hierarchy ends up being more natural.
The hope is that the mapping to URIs would be bijective. I wonder if there are some unexpected edge cases.
Another open question is how to also make it work with other schemes, for example URNs like urn:uuid:123e4567-e89b-12d3-a456-426614174000 or urn:oasis:names:specification:docbook:dtd:xml:4.1.2.> URIs complect identification, location and transport (http vs https distinction).
Practically speaking, it's considered best practice to only use http: schemes for URIs. If the URI is being serviced as a retrievable resource, it should get a 303. "See Other" is considered the most appropriate, since the data retrieved is generally not the actual resource. (e.g. the data representing a person is a not an actual human being)
> By document scoping you mean the # separator, right? If I understand correctly that is another inconsistency that bothers me
No. I'm referring to the fact that a graphs are defined by documents, like Turtle or RDF-XML documents. However, that's more precisely "were defined" since there are now a few syntaxes which allow multiple graphs in a document (e.g. TriG and TriX). Also, it's technically possible to change a prefix mapping during a document, but you'd almost never do that unless you're copy/pasting data from other sources into a single document.
The # separator is another issue again. Prefixes typically end with either # or /, but that's not required. Fortunately, I haven't seen anyone obtuse enough to use something different. π
There is a "best practice" around this. If prefix is used for everything in a defined group of URIs, then the # would be used. That's because the entire group may (and often will) appear in a single document, where the document is to be found at the URI of the schema. This is almost always the case for schemas and ontologies:
β’ rdf: http://www.w3.org/1999/02/22-rdf-syntax-ns#
β’ rdfs: http://www.w3.org/2000/01/rdf-schema#
β’ owl: http://www.w3.org/2002/07/owl#
β’ xsd: http://www.w3.org/2001/XMLSchema#
If you curl these addresses you'll get a document containing the full schema or ontology for the classes and properties in those specs.
Conversely, the / is used for resources that are theoretically unbounded. Catalogs, object identifiers, references for parts, etc.
This isn't quite what the https://www.w3.org/TR/cooluris/, but it's close and it seems to be how things settled out.
> By document scoping you mean the # separator, right? If I understand correctly that is another inconsistency that bothers me> No. I'm referring to the fact that a graphs are defined by documents, like Turtle or RDF-XML documents. However, that's more precisely "were defined" since there are now a few syntaxes which allow multiple graphs in a document (e.g. TriG and TriX). Also, it's technically possible to change a prefix mapping during a document, but you'd almost never do that unless you're copy/pasting data from other sources into a single document.
Named graphs are something I am planning to look into since merging graph data from multiple sources is the main reason for exploring RDF. My hope is that once having those mapped to RDF, then the merging process will become trivial just doing set union of triples and then adding some relation triples or rules to tie them together.
I saw your Raphael library, which is really nice for working with Turtle, has a todo to support TriG. But for now I use Turtle mostly for testing and the actual data is going to come from JSON/XML based sources, so I don't know yet if TriG is something that would be useful for that.> The # separator is another issue again. Prefixes typically end with either # or /, but that's not required. Fortunately, I haven't seen anyone obtuse enough to use something different. π
> There is a "best practice" around this. If prefix is used for everything in a defined group of URIs, then the # would be used. That's because the entire group may (and often will) appear in a single document, where the document is to be found at the URI of the schema. This is almost always the case for schemas and ontologies:
Yeah, it is my understanding that prefix does not have any defined meaning regarding namespacing, it is just a string concatenation. Then people by convention use / and # for namespacing purposes. Having definitions tied to document in general does not seem that great to me, because it is really hard to ensure any consistency. For example I've read an article where https://copia.posthaven.com/linked-data-and-overselling-the-http-uri-sche tool because they decide to change what they serve.
In other words would it be fair to formulate a rule of thumb that # separator is used for schemas/ontologies and / for data? Or is that too simplistic?
Sorryβ¦ I was working on this for a bit before you typed, so pretend that it appears immediately after my last messageβ¦
On your request to use Clojure namespaces, I note that while you're looking for a 2-way mapping between the prefix "namespace" (where "namespace" refers to the W3C definition) and Clojure namespaces, you still need to encode/decode these namespaces to map between the two. For instance, with your examples of [org.example.$ :as-alias ex] and the URI of ::ex/john, you get a keyword of: :org.example.$/john which needs to be decoded into the URI . The RDF namespace gets even messier with things like rdf:type β‘ :org.w3.www.$.1999.02.22-rdf-syntax-ns.$$/type
What if you had a data namespace, with something like:
(ns data)
(def registry (atom {}))
(defmacro prefix [pref namespace]
(let [local-data 'data
pname (name `~pref)
psym (symbol pname)
nsname (symbol (str "data." pname))]
`(do
(swap! data/registry assoc (str "data." ~pname) ~namespace)
(ns ~local-data (:require [~nsname :as-alias ~psym])))))
(defn decode [kw] (str (get @registry (namespace kw)) (name kw)))
Actually the way https://github.com/ont-app/vocabulary works is you define a namespace with metadata.
(ns my.example
{:vann/preferredNamespacePrefix "ex"
:vann/preferredNamespaceUri " "}
(:require
[ont-app.vocabulary.core :as voc])
...
)
Requiring the voc module would import namespaces and metadata for foaf and rdf
Then your triple woud be
> [:ex/john :rdf/type :foaf/Person]
You could also do things like...
> (voc/as-uri-string :ex/john)
""
> (voc/as-qname :ex/john)
"ex:john"
> (voc/as-uri-string "rdf:type")
""
> (voc/as-kwi "")
:rdfs/label
I like this one from @eric.d.scott. I hadn't used that module before. What I just wrote was very ad hoc.
Also, @kloud I think that's a reasonable rule-of-thumb. It's not strictly adhered to, but most vocabularies tend to do it that way
Going back a message, your comment on named graphs: > Named graphs are something I am planning to look into since merging graph data from multiple sources is the main reason for exploring RDF. My hope is that once having those mapped to RDF, then the merging process will become trivial just doing set union of triples and then adding some relation triples or rules to tie them together. I wasn't specifically looking at this, though it does overlap a little. Back in RDF 1.0, a "Graph" was formally defined by a "document". They had no name, and the document may not have a URI (as it could have been some kind of inline context). Even if a document did have a URI, there was no formal way to connect that to the graph. Then SPARQL 1.0 came along. This came out of a committee that was mostly individuals who were associated with the main RDF databases (Sesame, Jena, and Tucana. Note: I worked on Tucana), and who all had their own approaches. All of us had named our graphs, and we queried graphs by name. Hence, named graphs were in SPARQL 1.0. However, RDF didn't name graphs, so SPARQL also has the default graph. RDF 1.1 recognized the practicality of naming graphs, so they were introduced here too. TriG and TriX were informal standards and had already recognized them.
The voc module also has facilities like #voc/lstr and #voc/dstr
> (str #voc/lstr "dog@en")
"dog"
> (voc/tag (short 1))
#voc/dstr "1^^xsd:short"
> (voc/untag #voc/dstr "1^^xsd:short")
1
>(type *1)
java.lang.Short
And there is stuff in there that will automatically add prefixes to your SPARQL or Turtle, and a method to mint identifiers.@eric.d.scott cool, I am going to give the voc library a closer look. On a first look my concern is that if two independent modules pick the same prefix it would lead to a conflict. But maybe it does not happen in practice.
There is a facility for handling that possiblity, discussed here: https://github.com/ont-app/vocabulary?tab=readme-ov-file#prefix-collisions
@quoll Re named graphs: I see, thanks a lot, this is really helpful background π That explains the chain one needs to navigate through get to the actual vocabulary, which is not in RDF datasets but a SPARQL Service Description: trig -> rdf11-concepts -> rdf11-datasets -> sparql11-service-description If I don't encounter any limitations, I might be able to just load the data into Datomic and skip all that π