SQLGraph: An Efficient Relational-Based Property Graph Store (2015) [pdf]
research.google.com
research.google.com
Also which RDB did they use? I saw mention of Berkeley DB, but can we infer their schema would perform well with stock Postgres or MySQL?
Ideally, there should be code so we can reproduce independently.
The paper references 3 relational DBs: DB2, Oracle, and PostgreSQL...
Most modern relational databases such as DB2, Oracle or
Postgresql have features to support both relational and
nonrelational storage within the same database engine,
making it possible to perform an empirical comparison of
the utility of relational versus non-relational storage
structures for property graphs.
Presumably they used PostgreSQL since it's the only open-source RDBMS referenced, and it supports CET translations, which they used.The Berkeley DB reference is regarding the Titan backend datastore they used. Presumably they chose BDB over one of Titan's distributed backeds like Cassandra or HBase since BDB is single server so it's a more apt comparison to the other DBs.
On the last page of the paper, Table 8 includes Gremlin->SQL CET translations, i.e. Common Table Expressions: http://www.postgresql.org/docs/current/static/queries-with.h... , https://en.wikipedia.org/wiki/Hierarchical_and_recursive_que... .
Note the paper uses TinkerPop/Gremlin 2 step names, and since then TinkerPop/Gremlin 3 has been released (http://tinkerpop.apache.org), which includes many enhancements so the current Gremlin 3 names may differ a little from the ones they reference in column one of the table.
P.S. If you want to see some of the new crazy shit you can do in Gremlin 3, check out Marko's talk and paper from GraphDay a few weeks ago:
Quantum Processes in Graph Computing (https://www.youtube.com/watch?v=qRoAInXxgtc)
"Quantum Walks with Gremlin" http://arxiv.org/pdf/1511.06278v1.pdf
Common Table Expressions
Be careful to watch how your chosen DB processes CTEs and what it means for performance in your queries.In postgres they are "optimisation fences", something which there is significant resistance to changing[1][2], where-as elsewhere (at least in SQL Server) the query planner and engine can optimise across CTE boundaries (pushing search predicates back into them where possible). This can make a massive difference to the performance of some queries, potentially turning a full index scan (or worse, multiple scans) into a small number of seeks.
As always: test with realistic data sizes/patterns to make sure things work as you are expecting them to.
[1] see http://blog.2ndquadrant.com/postgresql-ctes-are-optimization...
[2] also https://news.ycombinator.com/item?id=7023907 for further discussion
SQL databases _could_ conceivably optimize foreign key constraints behind the scenes by replacing/augmenting the referencing column with the physical location of the foreign row, but what do you do when the target row moves around in the physical store? You'd have to hunt down all the references and replace them. If anything can point to anything, that could have a cascading effect.
You could have conceivably have an indirection here: Instead of pointing to the physical location of another row, you pointed to the physical location of a special intermediate file that itself could contain the physical location. I don't know of any database that uses any tricks like these.
You can, it just requires one further level of pointer indirection. Rather than pointing to the physical location of the data, you store a pointer to an offset in an array of pointers.
Alternatively you can point directly if you're prepared to update those pointers when you move things around. Luckily a graph database is ideally suited for this kind of operation because it makes discovering those pointers exceptionally simple and efficient.
Never met a vendor benchmark that wasn't faulty in some way... http://maxdemarzi.com/2015/10/16/benchmarks-and-supercharger...
[1] https://github.com/thinkaurelius/titan/releases/tag/0.4.0
SIGMOD is not like a dev conference where you can change the paper the day before your talk.