Does Neo4j have a better importing story now? I see a blog post from 2014 [1] that makes importing merely a million edges between 500 nodes sound like it's still a terribly difficult operation, giving me the impression that graph databases aren't quite ready for <s>big</s> medium data yet.
[1] http://jexp.de/blog/2014/06/load-csv-into-neo4j-quickly-and-...
If, for example, I wanted to load an N-Triples file that's approximately the size and shape of DBPedia, can I reasonably do so? What tools should I use to get the job done quickly without descending into a Java nightmare?
For the bulk of our work, we do what you'd expect -- load terabytes into HDFS, (Py)Spark for straight SQL and some join helper functions, and occasionally, add some GraphX scala libraries.
I'm curious -- what did you end up doing, & for what?
[1] http://conceptnet5.media.mit.edu
It's turned out to be a good input for machine learning about semantics, which has changed the goals of its representation a bit -- not only do I need to be able to load in data easily, I also need to be able to iterate over all of it. But some graph operations would be nice to have, too.
Many technical people I describe the project to immediately ask me what graph database I'm using, both before and after the ill-fated semester of grad school where I actually tried to use graph databases.
The answer to what I use now is: a bit of SQLite and some flat files. No need for HDFS, it still fits easily on a hard disk.
Disclaimer: I'm one of the developers of Stardog.
Stardog doesn't even show me a price for putting a billion edges into it, just an e-mail link marked "INQUIRE", so I have to assume it would be very, very expensive.
FWIW Developer/Enterprise versions are free to try and Community doesn't expire.
Thanks for the offer, but I'd only go with that model if there were no other options, and right now you're competing with SQLite and a filesystem, which have no additional costs.
The use-case you describe (a couple million nodes + edges) sounds like a fairly reasonable task for Neo4j.
If Neo4j has been a bad experience for you, I also recommend looking into TitanDB v1+, which scales horizontally (backed by Cassandra), although the query language (Gremlin) is not as easy to learn as Cypher.
I have a centralized logging system, but I can't imagine being so confused about it that I need my logs to tell me that two components interact with each other.
What don't I understand?
The moment more than a few people are involved with your systems, it can get so complex that visualizing dependencies can be extremely helpful and bring a lot of insights.
It's a really fascinating data problem, so we've been loving building tools for seeing into it!