Do your SPARQL queries turn into anything within orders of magnitude of an efficiently indexed SQL query when they get to the database?
Do your SPARQL queries turn into anything within orders of magnitude of an efficiently indexed SQL query when they get to the database?
I can't speak for other triple-stores as Jena is what we use, but I can say that comparing SPARQL queries against Jena TDB to SQL queries on MySQL or Postgres is like comparing apples and kangaroos.
In the sense that apples rot and Kangraroos are amazing animals which can do just about anything?
Because otherwise the comparison seems completely apt.
Jena works for RDF data. But the OP is correct in their broader point that RDF is rarely a good choice and SPARQL is a pretty horrible solution for querying it.
Note that the reply someone is about to write ("But RDF is a generalized self descriptive data model") means it is intended to solve the exact problem that a RDBMS+SQL solves. And if you add the "standardized fields" thing it also matches the REST interface+RDBMS+SQL comparison the OP made.
SPARQL is an excellent choice for querying RDF data (SQL is usable but awkward for querying EAV structured data).
> Note that the reply someone is about to write ("But RDF is a generalized self descriptive data model") means it is intended to solve the exact problem that a RDBMS+SQL solves.
RDF/EAV is more graph than relational structured. It doesn't solve the exact same problem.
Yeah 100% this. Comparing a graph data model and a relational data model while obviously possible isn't really all that fruitful so long as each is being use to solve the problem that they're the best fit for.
Firstly: RDF is an inadequate expression of most graphs, and SPARQL is a bad way to query graphs. See my comment here on this: https://news.ycombinator.com/item?id=14603090
Secondly: graphs storage is something which is very tempting in theory but very hard to get right in practice. I'm not going to say it is never appropriate (that is clearly untrue), but for most production applications it isn't the right choice.
I'd note for example that most social applications use a RDMS to store a single layer of friends (and then perhaps have a second graph DB for batch/stream processing of graph functions).
It's really not!
There's a reason why Tinkerpop is what most graph databases standardize on, and why things like Neo4J, DGraph, Caley, TitanDB etc (ie, all the graph DBs which people use when they want to build something and not "do the semantic web") don't use SPARQL.
Tinkerpop is a Java API, not an independent query language, and RDF data (while it is a way of modeling a graph) is not the model of most graph databases (it's a lower-level model than most graph databases expose, and is about as far from them as it is from the table model of SQL databases.)
An optimal Java API for graph databases with a more typical model is not an optimal query language for RDF data, for pretty much the same reason SQL isn't.
Now, what you describe is probably a sign may that RDF isn't the right exposed data model for many use cases (EAV-style representations are often used for deep internals, but there's probably a reason that outside of RDF most systems which use the model internally expose something more similar to the conventional relational or graph model to application developers), not that SPARQL isn't the right language for querying RDF data.
Gremlin is very widely supported across graph databases.
what you describe is probably a sign may that RDF isn't the right exposed data model for many use cases
Well. that's exactly what my claim is, so that's good!
not that SPARQL isn't the right language for querying RDF data.
Have you ever tried one of the alternatives? Try GraphQL on DGraph (Or Gizmo/Gremlin on Cayley) against a Freebase or DBPedia import. That's exactly the equivalent of SPARQL against RDF, and it's so much better.
The standardization in REST is basically "read stuff" and "write stuff". For everything else you're supposed to look up the API docs and write a specialised client.
I can query a SPARQL endpoint for a list of people and their friends, sorted by age - without knowing anything about that endpoint. I can also merge results from different endpoints without worrying whether they use slightly different data formats.
> I can query a SPARQL endpoint for a list of people and their friends, sorted by age - without knowing anything about that endpoint.
No you can't. I am certain that you can't.
To do this, you would need:
- A social networking service that uses SPARQL
- People to actually use that service
- Knowing the schema that would represent things like "friend" and "age"
- A model of permissions that indicates that somehow you're allowed to know the age of people's friends (seriously, how are you allowed to know this, that's creepy)
- A way to express that permission alongside your SPARQL query, which probably means you need to expand the query to include a representation of your identity and permissions
- A way for the SPARQL endpoint to authenticate that you have that permission (you will definitely need to look up API docs for this, as it will involve sending some sort of crypto token out-of-band)
- A container format for the RDF responses you get that can express things like "you don't have permission for that query"
Then please point me to that standard. In RDF, that woukd be FOAF for example.
> A social networking service that uses SPARQL
- People to actually use that service
Yes, for querying an endpoint, I need an endpoint. No way.
My point is that even if I have such an endpoint as a REST API, I can't directly go on to query it because I'll first have to write a specific client tailored to its API and data model first, then think about how I convert it into my own. If I want to match up accounts from Facebook, Twitter and Mom-and-Pop-BBS, I'll have to deal with three different APIs and three different data models. If those sites provided SPARQL endpoints, I'd only have one of them.
> Knowing the schema that would represent things like "friend" and "age"
Defined by FOAF, see above.
> A model of permissions that indicates that somehow you're allowed to know the age of people's friends (seriously, how are you allowed to know this)
That's the responsibility of the endpoint, not mine. I don't see why that would be a hard problem (I figure you'd define permissions on different RDF properties and types) but I admit I don't know much about it.
>A way to express that permission in your SPARQL query
I send my (authenticated) query and if I don't have sufficient permissions, the server will hopefully return "nope". Why would I need to send more?
Yes, some sort of authentication is obviously needed, but there are enough standards to use for that (any sort if HTTP auth method, OAuth, OpenID etc)
Note my point wasn't that I can query endpoint X out of the blue and expect to get all the data - but that I don't have to write specific code to deal with endpoint X. Obviously I have to get permission somehow, but ideally, the only endpoint-specific thing I have to do is to fill out a registration form.
Depending on the use-case, you might not even need auth at all if your endpoint is restricted. We also have authless, restricted REST endpoints today that seem to work well: They're called web pages.
It worked, sort of. But only after I mapped the many different representations of "age" used by the different end points.
I don't remember the specifics, but even in DBPedia alone you have to deal with the properties and the ontology namespace. Then YAGO uses that but brings in other sources and puts them in their own fields. Freebase does (did) its own things.. etc etc.
It was a long, long way from the "you don't need to know anything" utopia you describe.
In summary, there really is no advantage over mapping from a REST endpoint.
Plus, the database endpoints are slow (and even worst - have high variance in performance). I ended up downloading the dumps and hosting them all myself because the servers were so slow and unreliable.
Probably more efficient, if it's an RDF- (or EAV) oriented datastore underneath, and merely similarly efficient if it's actually layered on top of SQL.