LemonGraph – A log-based transactional graph
github.com
github.com
...actually this has been done [2] but the project looks abandoned. So NSA guys you should get right on this.
https://motherboard.vice.com/en_us/article/d737mx/the-fbi-ca...
EDIT: oops that was 2014, apparently the FBI isn't having that issue anymore
https://motherboard.vice.com/en_us/article/aepj4p/fbi-mariju...
Huge amount of stuff devoted to managing ETL.
The link mentions that its use case is streaming seed set expansion, which allows you to identify communities based on a set of seeds. I wrote more about that in this comment: https://news.ycombinator.com/item?id=17335873
I actually love the idea of a memory-mapped database, I've often thought memory mapping isn't taken advantage of enough.
"..log-based transactional graph (nodes/edges/properties).. ..primary use case is to support streaming seed set expansion." -- I'm totally lost.
I know these kind of software is targeted at developers, but it won't hurt to give analogy like "Uber for XXX" like in startup pitches. e.g. "It's like <put popular product name here e.g. MySQL> but <differentiating factors>".
Streaming is just referring to its ability to be constantly updated by incoming data.
So it's like a database but for a much more narrow use case and can perform much better in those cases.
There are a bunch of them on the market, Neo4j is probably the most popular (and has lots of good quality introductory text on the website and on youtube). Graph databases are key to many major internet services, eg Google, Twitter, and Facebook are all just really big graph databases.
This particular graph database stores all its data in a single file trading speed and simplicity off against flexibility. I'm not an expert but it seems like it would work very well for search queries, but poorly for tasks involving a lot of contributors like a chat server.
The particular use case mentioned for LemonGraph is "streaming seed set expansion", which, given a set of "seeds" can expand that set based on their communication patterns to find people that are likely in the same community or overlapping communities.
E.g., if you have a few known terrorists and a database of metadata about phone calls or internet communications (emails etc.), you can analyze who your known subjects (seeds) talk to and who those people talk to, etc., to identify communities. This relies on the fact that there tends to be cross-communication between people in the same community.
This kind of analysis can often reveal the structure of communities, like where the headquarters, who the boss is, etc.
In industry and law enforcement, the same kinds of approaches can be used to identify fraud of various kinds.
"What it's like" is a graph database, in this case based on a fast in-memory database to support high speed graph analysis on a single machine.