High-Performance Graph Databases
arxiv.org
arxiv.org
Those are related but are distinct from each other.
And sure about 95% of companies would have their needs met with a simpler system but that does leave a lot of companies who will not. And for those of us in say finance doing customer/fraud analytics I would welcome all the performance I can get.
The paper has "Scale to Hundreds of Thousands of Cores" in the title. I have not yet read the paper but it seems unlikely it doesn't talk about scalability.
You can have slow queries with 10GB of data just like you can have fast queries with 10PB of data.
This is really common across article-comment platforms; is anyone interested in discussing how to incentivise comment sections that have read the paper?
In these kinds of workloads you quickly run into performance bottlenecks. Even in-memory analyses need care to avoid conplete pointer chasing slowdowns.
I do still hope this is fast in like a single CPU 32 core 64GB system with an SSD. But if this takes a cluster to be useful, then I will still love it.
Yeah, but the hacker fascination is what drives progress. You could have made the same type of argument about ML, and we would have been content with MNIST.
One of the simpler supported backends for our Modality product (https://auxon.io/products/modality), which results in a data model that’s a special case of a DAG for modeling big piles of casually correlated events from piles and piles of distributed components for “system of systems” use cases, is built using SQLite, and the scaling limiter is almost always how efficiently the traces & telemetry can be exfiltrated from the systems under test/observation before how fast the ingest path can actually record things becomes a problem.
That said, I do love me some RDMA action. 10 years ago I was fiddling with getting Erlang clustering working via RDMA on a little 5 node Infiniband cluster. To mixed results.
- proportion of jobs (not companies) requiring extreme scale - the fact that non extreme scales are the long tail doesn't mean it's a fat tail
- proportion of buyers/potential users that walk away from the inability to handle extreme scale
... and more sarcastically
- proportion of articles about extreme scale
- proportion of repos about extreme scale
> Being able to constructively challenge assumptions is an incredibly valuable job skill
I would agree but ...
> good managers
... are few and far between.
Had one boss get mad that I reduced the database footprint by 94% - why? Because he wrote the initial implementation and refused to believe that his baby, which cost so much space because of how awesome it was, could fit into 5GB.
But challenging the status quo has gotten me to where I am, so I wont stop it anytime soon :)
Definitely a data heavy product, wherever it is that you're offering.
(Unless you keep large blobs in the DB. But database scale has more to do with records than raw storage.)
And these are seemingly huge services.
And yet…
Aside from coordinating big groups to write tons of papers, he does a bunch of impressive wilderness exploration. I recommend checking out his website, there's some stunning photos:
I know a matrix multiplication can be a breadth first search if you use a 1 for each column you want to breadth search.
How do you shard a graph data structure?
subject -> predicate -> object;
Example: `<bob> <isa> <person>` (from this RDF primer [1]).
Sharding data that follows this model sounds easy? Any sort of rule and/or hashing over the triples could be used to send triples to arbitrary servers.
I created a D3 visualiser from Jena Fuseki database
https://camo.githubusercontent.com/3064a94d00812c1373c4eb3b2...
It renders RDF/N3 relationships. Unfortunately the code is trapped in an Sqlite jsbin file :-)
It never occurred to me that it would be that trivial to shard, you could mapreduce the queries. I just worry about how you would do joins on data that live on independent nodes.
If a value is on server A and another on server B and there would be a relationship between them, how would you "join" this data unless you colocated join keys.
I suppose that it is possible to do something like this. We've explored the idea of synthesizing a workload that behaves like an application's real workload. The idea was to do offline testing with the synthesized workload to find optimal configurations that can then be applied to the production database. Our results were inconclusive (not enough real workloads to verify). But the idea could be extended for what you are proposing.
But in terms of relational vs graph, my opinions about the matter are public:
https://www.theregister.com/Debates/2023/03/06/great_graph_d...
You're looking for patterns across large numbers of entities and relationships.
And ideally you want this all done in real-time so you can stop transactions before they are approved.
Generally you could say that application logic/queries being more concerned about links or linked entity attributes than direct entity attribute values is a sign that a graph DB may be a good fit. I'd say about 50% of db using apps are like this. You can get an idea by looking at their benchmark workload descriptions.
If you mean fully RDMA offloaded huge cluster DB systems, then I don't know the answer but suspect this kind of thing hasn't made it to industry yet:
"GDA internally uses a distributed hashtable (DHT) to resolve dif- ferent performance-critical tasks conducted under the hood, such as mapping application vertex IDs to internal GDA IDs. For highest performance and scalability, GDA’s DHT is fully-offloaded, i.e., it only uses one-sided communication, implemented with RDMA puts, gets, atomics, and flushes. Its design is lock-free, it incorporates sharding, and it uses distributed chaining for collision resolution. To the best of our knowledge, this is the first DHT with all its oper- ations being fully offloaded, including deletes. The DHT consists of a table (to store the buckets) and a heap (to store linked lists for chained elements)."
- Feature generation for machine learning model training (particularly popular with fraud detection at financial institutions) - social networks (think LinkedIn or Facebook) - supply chain analysis and optimization - healthcare patient data analysis (looking at similar patients to recommend treatments or do large scale analysis) - user identification (eg taking lots of data points and tying to a specific user). There’s a more specific name for this I can’t remember off the top of my head.