I'm still not sure if this is the right road to go down. Just from reading papers it's hard to tell how much of this is hype and over-engineering and how much is solid. Anyone have some RDF/OWL/Jens/SPARQL stories they want to share?
I'm still not sure if this is the right road to go down. Just from reading papers it's hard to tell how much of this is hype and over-engineering and how much is solid. Anyone have some RDF/OWL/Jens/SPARQL stories they want to share?
Once I had "everything" in RDF, using SPARQL queries to nose around was a really great way to find arbitrary relationships. I then used GraphViz to map out everything and it was very helpful to see how everything fit together. Could've probably used a regular database, or a key value store, but working with just flat files transformations at the command line was nice.
I've also since used it manage a knowledge base and generate documentation and presentations. I haven't even used much of OWL or higher levels of abstraction yet, but just from a hack up something standpoint I think it is a pretty nice set of ideas to get basic graph operations just about anywhere you may need it, in the shell, in the browser, or in a backend project. Large sets of data may require something else, but anything less than 1g of data is rather performant IMHO.
To get a good start with RDF etc, I recommend 'Semantic Web for the Working Ontologist' (second edition or later!) [3]. It explains how inferencing works by doing inferencing with SPARQL queries.
RDF is a powerful technology, but takes time to get acquainted with.
What I dislike is that there's not many libraries (that I know) for working with it. I started a Rust library (Rome) for working with Turtle files and would like to have time to make it into a SPARQL engine.
[0] http://vocbench.uniroma2.it/ [1] https://github.com/UKGovLD/registry-core [2] http://www.pilod.nl/wiki/Platform_Linked_Data_Nederland [3] http://workingontologist.org/
Refinitiv (formerly Thomson Reuters Financial and Risk) knowledge graph is built completely on the RDF stack: https://www.refinitiv.com/en/products/knowledge-graph-feed
When I talked to them in late 2017, they told me they have 100 billion triples in their database, plus more in a versioning back-end. Their triplestore is open source: https://github.com/CM-Well/CM-Well
Several government-agencies all over the world start to build public RDF knowledge graphs. I'm closely involved in the one from the Swiss government, see my presentation from last week http://presentations.zazuko.com/Swiss-LOD-Platform/
There are similar projects in other countries like the Netherlands, Belgium, UK, etc. This stack makes a lot of sense for open data, as you can do some pretty crazy queries without spending 2 days on preparing your data. See for example the Swiss Open Data Advent Calendar of 2018: https://twitter.com/linkedktk/status/1076064066525949952
As I said there are many "behind the firewall" use-cases where people use the stack exactly because of its features like OWL. Yes it comes at a price (bootstrapping is not really super easy) but this is stuff we will still run in 40 years from now. I see it in:
Finance: Fraud detection, compliance, customer 360° views, ... Stardog (https://www.stardog.com/) lists Moody's, BNY Mellon and National Bank of Canada as customers, last week I've met someone from Credit Suisse which is Mr. RDF there. * Production: You have a ton of databases containing products you create but there is no way to figure out what a final product consists of as the data is scattered across at least 5 of them. The automotive supplier I talk about here is using RDF to get that view.
Life sciences: The largest RDF dataset available to the public is UniProt and related datasets. In total they provide a SPARQL endpoint (RDF database) with 50 billion (!!) triples available. This is a highly popular dataset and is used in pretty much every larger pharmaceutical enterprise as well. See https://www.uniprot.org/ as a starting point. I know at least of one large life sciences company that just recently decided that RDF will be the base of all future data unification standards within the organization.
Insurance business: One of our customers is using RDF to unify a ton of different systems and get the 360° view as well about their customers.
RDF is an absolutely amazing stack and I do not see anything else available that gets remotely close to the power of it. The day I find something more powerful, I will be the first using it. But most of the time people dismissing RDF have zero clue about what it really can do.
If I have transcribed voice convo data, with date/time, names, location, sentiment and extracted subject matter keywords, would Jena/RDF and/or related tools be appropriate for exploring relationships and trends between data points?
Thank you.
For us, RDF seems like the only technology that can easily adapt to the large number of data types that we envision collecting.