I'm curious how does it perform in comparison of neo4j. Especially in memory usage.
I'm curious how does it perform in comparison of neo4j. Especially in memory usage.
[0]: http://data.gdeltproject.org/documentation/GDELT-Global_Know... https://blog.gdeltproject.org/gdelt-2-0-our-global-world-in-...
Edit: I just read some more, and it seem like this is a graph-native "brother" of cockroachdb. Which was the one I previously cosndidered, until learning of the inadequate down-pushing ability of graph queries, which would increase network needs a lot. Let me know if you want me to try and notify you once there are benchmarks or production tests successfully completed.
2-digit TB of graph data is for sure a serious graph :) Depending on what you actually want to analyze and how the graph is structured, you can either use our Pregel integration for high-level analytics of the data set or use SmartGraphs to shard the data to a cluster and perform in-depth traversals, pattern matching and such stuff. SmartGraphs is an Enterprise feature but free for evaluation and testing. Please note that the suitability of SmartGraphs depends on the structure of the graph itself (e.g. do you have or can identify communities to shard by efficiently)
We created a series of tutorials which guide you through such a process https://www.arangodb.com/pregel-community-detection/ (next step always at the end of each tutorial). You can choose between two storage engines. Think RocksDB is the better choice in your case, as everything is persisted to disk (data/indexes) and you can configure how much main memory should be used, so you can configure the trade-off between performance and main memory yourself. If you have fast SSDs it's even better, as RocksDB is optimized for that.
https://www.elastic.co/guide/en/elasticsearch/reference/curr...
As an example, I have the HN data set on my laptop - about 15 GB of data and I have the max mem set to 3GB (heap is usually ~1) and I can search, aggregate, etc. very quickly without memory problems.
I've never tested the memory usage difference between the two, but it should be an interesting test to perform.
Generally speaking it seems like a common practice now days to load large portions of your data from disk to memory for performance reasons obviously, no one is willing to pay the penalty hit of disk seek when it comes to real time data querying.