The "semantic web" or "open linked data" concepts never really took off the way people had hoped, but there's till a ton of utility in the underlying standards so you'll tend to find it wherever you need complex, flexible schemas that with good interoperability between different entities.
These so-called "semantic web" technologies seem to come into their own when there's large scale organizations interfacing without a common reference frame. Like one org that does a spec from a programmer standpoint, and another org does one from a formal linguistics standpoint, then they have to integrate. For example, the USDoD Logistics steering group makes a spec for parts data from their requirements based on MTTF, cost, sparing, shelved space. USN makes a spec for parts data based on burn rate, transport, fuel type. It goes on and on like this, repeat a few dozen times, and you have a dump truck full of specs doing the same thing. See where I'm going here? They're speccing out the same thing from their own ivory towers, and - here's the kicker for those trying to LLM their way out of the situation - none of them are going to show their data to anyone else. The only thing that's exposed is the semantics. ARTT/CredEng is - or was, I am not sure if the program OR CredReg is still healthy - trying to solve this by unifying the semantics.
Ultimately someone's got to come along and give all these people a kick in the pants, one way or the other. You can't just float a boat around the ocean with no missiles, not these days.
https://data.europa.eu/data/sparql?locale=en#
I did struggle to find what I wanted. It's a labyrinth of metadata, and I was looking for the structured regulation text itself. In the end I stuck with good old fashioned XHTML scraping
I have seen many more useful tools come out of LLM in the short time it has been available than the entire 10 years working with academics using RDF, Ontologies etc. RDF is too difficult to use and has inadequate tooling. LLM is only going to get better.
There's not much application for knowledge graphs in e.g. a CRUD app of customer names and addresses, but turns out there are an unlimited number of things you can describe about e.g. a protein, and you can't just design one schema because you don't know how it's going to be queried.
See: https://bioregistry.io for countless examples of public datasets used everywhere from academia to "big pharma".
https://ontology2.com/essays/LookingForMetadataInAllTheWrong...
I am not well versed in the other RDF technologies. I haven't paid any attention to ontologies, or OWL or any of that stuff. I just use raw RDF, and defined my own vocabularies for everything, including structure. For example, I have my own type property. RDF also has one, but I just made my own. I have my own structure system to mostly bring order to how things are displayed, or created, etc. I am pretty much as far from the semantic web as you can get.
Since everything is "just a triple" it makes it easy to share data. So, it'll be straight forward to import and export artifacts out of my system and share them with others.
And I get SPARQL "for free", so even after new data structures are added, they're still queried like the first class ones the tool already knows about. SPARQL is pretty neat.
At the moment, I have a mostly complete RDF CRUD tool, with some first class interface prototypes (by first class I mean I have forms and UI specifically for those data types, rather than a generic resource form), and really like working with it. My DB has about 3.5M triples in it.
> I just use raw RDF, and defined my own vocabularies for everything, including structure.
I think this is best approach. The ontology part was more of a hindrance for me way back in 2010's when I was experimenting with semantic web technologies (using dbpedia as a source of my data) and I really hard tried to avoid going of the beaten path (no matter how flaky it seemed) as a junior level developer.
Performance? I have nothing to compare it too. I can't complain. I know whenever I saw info about triple stores in the past, they only seem to crow about was how fast it takes to import things. "Eleventy trillion triples per bleem!" I guess nobody ever actually queries the data, they just store it.
I routinely export the model to a file, and that takes seconds (<10), using the N3 format (RDF/XML takes a very long time). I export the model to make sure my changes are reflecting properly. The resulting file is 176MB. If I read that into a new TDB instance, I can load it in 25s. Since I've been importing from a SQLite master, the resulting TDB data is roughly the same size as the SQLite DB file. My import from SQLite takes longer than 25s, and that's just data shoving from the SQL data to the RDF. I'm sure I'm the bottleneck in that case, I probably commit too much for one thing.
As for queries, well, it either can find it or it can't. It's either trivially indexed (I assume each of the properties of the triple are indexed), or it table scans. Internally when you do a query from Java, you basically set the base net you want to throw (you've only got 3 values to work with) and iterate through to filter it. When I did my "select count" query to count all the triples, that took a beat or two to be sure as it hoovered the entirety of the model and cursored through it. I have not done any crazy SPARQL queries (I can barely spell SPARQL), so I don't know what kind of decisions is makes, but, in the end, there's really only a few ways you can actually query a triple store.
Now, I have a recent Intel iMac I'm running this on, so that may well impact things as well. I have no idea how much memory I'm using, it hasn't been a problem. I've done no tuning whatsoever, I honestly don't know what tuning is available.
I do not foresee my dataset growing much more, so TDB is "fast enough" for my purposes. All told I'm pretty happy with everything.
RDF is great for annotating protein interactions, for example.
https://developers.google.com/search/docs/appearance/structu...
TBL's vision for knowledge graphs is even older than the web. But should it be W3C's job to invent new tech? Does W3C's track record, legal and financial standing invite further standardization work? Their HTML and SVG charters have basically ceased working and W3C's last (final?) HTML recommendation is based on WHATWG HTML Review Draft January, 2020 [1].
[1]: https://data.gov/ [2]: https://data.europa.eu/
[1] https://www.w3.org/TR/vocab-dcat-3/
[2] https://joinup.ec.europa.eu/collection/semantic-interoperabi...
[3] e.g. https://www.dcat-ap.de/ or https://docs.dataportal.se/dcat/en/
Trivial (as in ‘cat graphx graphy’) merging of complex data graphs is just too powerful.
If the need is in fact communicating $100bn design specifications between multiple transnational engineering and construction co’s, even the most grounded and pragmatic engineer will gladly inplement RDF and ontologies in the hot path. It’s the best tool for the job on the merits.
Sometimes they need a little nudge. A python or C# library here or there, just to take the edge off, you know?
Was also told by some of my colleagues that they were earning a lot consulting in the medical world with semantic data.
https://newsroom.accenture.com/news/accenture-invests-in-sta...