(Note: I've actually written a graph database from scratch, for exactly these reasons.)
(Note: I've actually written a graph database from scratch, for exactly these reasons.)
Fixed that for you. ;-)
Actually a few RDMBSes (Oracle, Postgres, MS SQL) have graph extensions. I've never used any of them, but I assume they work around some of the basic unsuitability of traditional tables for storing graphs.
Also, in this case, "probably less" is "multiple orders of magnitude slower".
const unsigned *edge_indexes = (const unsigned *) mmaped_structure_on_disk[disk_offsets[index]];
That's not the same as just being able to access those via an API; locality in this case means that you shouldn't need extra seeks for every single value, nor have to make a bunch of round trips through SQL.That is the primary difference between traditional relational databases and column-oriented databases. Normal relational databases have row-based locality; column-oriented databases have column-based locality.
So the question is, do we actually want to run graph algorithms or is the data graph structured for other reasons? You're implying that choosing a graph representation means we want to perform graph analyses. I disagree with that if we're still talking about the semantic web.
RDF is a general purpose knowledge representation model. The triple structure lends itself well to combining data from different sources with little coordination. It happens to form a graph, but running graph algorithms is just one of many special purpose problems.
I'm primarily replying to:
> What is a graph database? A miserable little pile of joins.
> Though to be serious: what do you expect a graph database to provide that sqlite cannot / does not do efficiently?
That seemed like a general questions of, "What are graph databases for, and why would someone use them?" And I'm trying to answer that question.
If there was some kind of "normal" query on thousands-to-millions of items that was prohibitively terrible on SQLite but not on graph-database-X, yea - I'm interested :) And I totally buy that graph DBs are better at graph queries in general. I just have yet to hit these kinds of limits in my use of SQLite (a fair number of instances with tens of gigabytes, a few with billions of rows) - a sprinkling of reasonable database design addresses almost all issues.
The main one I can see is that, with longer-term use, SQLite's lack of any way to force locality would be fairly crippling. You'd need to make a reasonable sort order and periodically re-insert data in that order to optimize / vacuum. That's... technically achievable, but is a big downside compared to something that can dynamically organize / optimize it based on [some heuristic].