Titan Graph Database Integration with DynamoDB
allthingsdistributed.com
allthingsdistributed.com
Back when DataStax acquired Aurelius, they announced [1] that they would stop developing Titan, and for a while it looked like it was completely dead. It seems DataStax are maintaining it somewhat, but there have been only 74 commits this year, all/almost of it from DataStax, and there's not a lot of meat on those commits. Plenty of open issues in the Github tracker.
For example, I noticed that Wikidata was considering Titan but dropped it [2] after the Aurelius announcement, and ended up with a fairly obscure database called BlazeGraph instead.
[1] https://groups.google.com/forum/#!topic/aureliusgraphs/c07WE...
[2] https://lists.wikimedia.org/pipermail/wikidata-tech/2015-Mar...
From your [1] we find the follow-up [1a] which clearly says work will not stop.
[1a]: https://groups.google.com/forum/#!topic/aureliusgraphs/WTNYY...
On the other hand the usecase of wikidata and uniprot are quite different from most graph database deployments who have only internal or controlled access via APIs.
Still UniProt is a graph with 3 billion nodes and 15 billion edges so not tiny but not humongous either. Wikidata is a bit smaller if I recall.
Titan seems to have a different use case, much more orientated to graph traversal than analytics on graph modelled data. So I can understand that many systems need something like titan.
You would use them if you want to experience the bleeding edge of databases and exploring uncharted territory excites you, and not if you want to get webapps built.
We've worked on it for over a decade, it's used in production by thousands of community users, hundreds of customers and 75+ Global 2000 companies (see http://neo4j.com/customers). For many of those Neo4j is used in business critical use cases, i.e. they require Neo4j to be up and running every minute of every day or it'll show up in their next earnings call. If you've shopped online or in a US retail store this week for example, it's very likely that you've used Neo4j. There's rich support for pretty much any programming language and framework out there, an ecosystem of consulting partners whose sole business it is to do Neo4j implementations, 10+ books written specifically about Neo4j, rich online training, formal enterprise support backed by a global commercial organization, an active community. What's missing?
I'm not trying to be facetious -- I'm genuinely curious as to what you feel is missing to consider it mature.
It had showstopping security problems when bound to anything but 127.0.0.1, so I came up with a software firewall to put around it and hoped for the best. It promised Lucene search but its implementation was full of Lucene injections, unless I escaped every special character I could think of like a freaking PHP programmer. There was no way to get data in faster than a slow trickle, unless that data was somehow already in another Neo4j database. Doing any interesting graph operations led to interesting messages about running out of "PermGen". And before I could even get all the data in, it had consumed enough resources to blow my academic AWS budget for months.
I was on the mailing list looking for support, and found it pretty lacking. The best I ever got was a bunch of Java code to try (my code is in Python).
I use SQLite now. It doesn't do very much, but it does what it's supposed to, and that's great.
If Neo4J has improved significantly since then, forgive me that I'm not rushing back to try it again.
And thanks for being specific (amazed that you remember specific issues from five years ago!). I don't remember the 127.0.0.1 security problems, but I don't hear anything about them so my guess is they've been addressed. We have a lot of finance and government customers that have high requirements on security. As for your Lucene issues, we did a complete overhaul of our search and indexing story in Neo4j 2.0 (released late 2013). We've continuously improved import performance (which has traditionally been a weak spot) and Neo4j 2.2 includes a batch importer which injects >1M records / sec sustained pace at scale (10s of billions of records) on commodity hardware. As for the memory management issues, we like many other data products written in Java struggled with GC for a long time, and like many others we ultimately concluded that we had to move a lot of the critical parts off heap / manage the memory ourselves, which significantly improved memory utilization.
I understand that you got stung historically and therefore hesitate to check us out again. And if SQLite is working well for you, there's no need to! But Neo4j and the graph space has matured a LOT since 2010 and fortunately I don't think your "bleeding edge" experience from 4-5 years ago will be replicated anymore for someone coming new into the space.
Thanks for the feedback.
While neo4j has it's proponents. The lack of standards support means that as a data provider it's hard to support.
Then for our end users, we would need to hack in a namespace convention to avoid issues when integrating our data.
Then TinkerPop misses the SERVICE concept for federated querying in SPARQL1.1, which is essential for our endusers who do knowledge discovery (i.e. small biology labs without the inhouse capability of running their own large databases).
If you're in SF or Oakland and I could buy you lunch or a beer and talk about graph DB's with you for an hour, please find my email in my profile and drop me a line :)
Then I decided to switch to Neo4j, and I was up and running in literally an afternoon. Also, its Java API is very well designed, which allowed me to fork and extend an integration plugin with Elastic in a couple days (https://github.com/jazzido/neo4j-elasticsearch/)
Anyway, I'm going to give titan a second look with dynamo. I'm using OrientDB and neo4j right now. Both have major scaling issues, we haven't reached their limits yet but expect too very soon. I am wary of Titan because of Tinkerpop.