An Introduction to the Resource Description Framework (1998)
dlib.org
dlib.org
I think Jena, a free and open source RDF framework, has been around for the last 20 years, and still releases pretty often. If you are considering using a graph database for your next project, you should give it a try! Here are some notes on how to get into RDF [1].
--
RDF is really at its core such a simple concept! What's built on top can go pretty deep though. RDF is a technology to describe graphs, with tools on top that allow querying, traversing, validating, and inferring information.
The way RDF describes graphs is with the use of a "triple". This triple has the form:
Subject -> Property -> Object
A sample triple could be: Tarantino Directed Kill-Bill.
The power of the model comes from the fact the elements of the triple use unique IDs by mean of URIs, so a more realistic triple could be:
<http://imdb.com/tt0266697> <http://movies.com/job/director> "Kill Bill Vol. 1"@en.
In this example, the object is a literal instead of a entity with its own URI: the English title of a movie. As long as someone else uses the same subject URI, I may automatically discover new things about my own data by merging someone else triples into mine. There's also an extension that allows naming the graph where the triple lives: the triple becomes a "quad". The name of the graph should be an URI: <https://me.com/my-movies>.
Through the years there's been more stuff built on top of the model: various serialization formats (most of them plain text based), a query language (SPARQL, similarities with Datalog), frameworks for "semantic inference" (OWL), schemata and constraints, asking the graph to be a certain shape (SPIN, SHEX, SHACL), an easy way to make statements about the triples, like "this triple was created on this date by this person" (RDF*), and I'm sure a lot more stuff I'm forgetting or I don't know about :-).
1: https://www.reddit.com/r/semanticweb/comments/fxtaxe/books_o...
I really liked the concept and we borrowed the "triple" approach in our next iterations, but we really didn't get any benefit from using RDF.
Quite possible it was our fault.
Wikidata is an RDF instance.
The largest chemical database in the world is based on RDF (PubChem).
There are hundred, thousands of other RDF instances.
As for what I'm doing, it's a project that is taking various siloed fan-created data on the web and collecting, aggregating, and transforming it into RDF for the purpose of being queried by researchers. I think RDF and RDF-adjacent things (e.g. knowledge organization, ontology application) are useful here, but the actual querying / data analysis requiring an RDF transformation is maybe questionable.
Wikipedia submissions are ok when the topic is obscure and there isn't any other good third party article to submit. They're bad when the topic is generic or well-covered. Wikipedia itself is already a generic site, and the combination of generic site and generic topic leads to low-quality discussion: https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que.... But when the topic is unpredictable and mostly unheard-of, such submissions can be quite good.
In this case a bit of googling revealed a 1998 article that hasn't been discussed on HN before, so I've changed the URL to that from https://en.wikipedia.org/wiki/Resource_Description_Framework. Any other good non-obvious article would have done as well; people will probably mostly comment on RDF and the semantic web in general.
One particularly memorable one was https://en.wikipedia.org/wiki/199_398_500_A (though I can no longer find the HN link to it); it's then easy enough for me to Google around for more context if I care to do so.