What's the advantage of using all these different things in one system? You can do all of this in datalog. You get strong eventual consistency naturally. LLMs know how to write it. It's type safe. JS implementations exist [0].
What's the advantage of using all these different things in one system? You can do all of this in datalog. You get strong eventual consistency naturally. LLMs know how to write it. It's type safe. JS implementations exist [0].
Zod/Valibot/ArkType/Standard Schema support because you need a way to define your schema and this allows for that at runtime and compile time.
Y.js as a backing store because I needed to support offline sync, branching/forking, and I use Y.js for collaborative editing in my product, so I needed to be able to store the various CRDT types as properties within the graph. e.g. you can have a `description` property on your vertices or edges that is backed by a Y.Text or Y.XmlElement
Cypher because until the arrival of codemode it wasn't feasible to have LLMs write queries using the Gremlin-like API and LLMs already know Cypher.
Most of all though, this was an experiment that ended up being useful.
I'm personally of the opinion that "graph databases" should be relational databases; the relational model can subsume "graph" queries, but for all the reasons Codd laid out back in the 60s... network (aka connected graph) databases cannot do the latter.
Let the query planner figure out the connectivity story, not a hardcoded data model.
% 1. Base case: Directly connected systems (1 hop) with bandwidth > 10
fast_path(StartSys, EndSys, 1) :-
link(StartSys, EndSys, Bandwidth),
Bandwidth > 10.
% 2. Recursive case: N-hop connections via an intermediate system
fast_path(StartSys, EndSys, Hops) :-
fast_path(StartSys, IntermediateSys, PrevHops),
link(IntermediateSys, EndSys, Bandwidth),
Bandwidth > 10,
Hops = PrevHops + 1.
% 3. The Query: Find all systems connected to 'System_A' within 5 hops
?- fast_path('System_A', TargetSystem, Hops), Hops <= 5.
or in RelationalAI's "Rel" language, such as I remember it, this is AI assisted it could be wrong: // 1. Base case: Directly connected systems (1 hop)
def fast_path(start_sys, end_sys, hops) =
exists(bw: link(start_sys, end_sys, bw) and bw > 10 and hops = 1)
// 2. Recursive case: Traverse to the next system
def fast_path(start_sys, end_sys, hops) =
exists(mid_sys, prev_hops, bw:
fast_path(start_sys, mid_sys, prev_hops) and
link(mid_sys, end_sys, bw) and bw > 10 and hops = prev_hops + 1)
// 3. The Query: Select targets connected to "System_A" within 5 hops
def output(target_sys, hops) =
fast_path("System_A", target_sys, hops) and hops <= 5
https://www.relational.ai/post/graph-normal-formhttps://www.dataversity.net/articles/say-hello-to-graph-norm...
...
That said, modern SQL can do this just fine, just... much harder to read.
WITH RECURSIVE fast_path AS (
-- 1. Base case: Directly connected systems from our starting node
SELECT
start_sys,
end_sys,
1 AS hops
FROM link
WHERE start_sys = 'System_A' AND bandwidth > 10
UNION ALL
-- 2. Recursive case: Traverse to the next system
SELECT
fp.start_sys,
l.end_sys,
fp.hops + 1
FROM fast_path fp
JOIN link l ON fp.end_sys = l.start_sys
WHERE l.bandwidth > 10 AND fp.hops < 5
)
-- 3. The Query: Select the generated graph paths
SELECT * FROM fast_path;RelationalAI's model is very cool but it is cloud only software.
Modern computers are also much more efficient at batch / vector processing than they are pointer hops. By either CPU or GPU, the whole system is much more tuned for working with data as sets / tensors rather than "hopping" through the data via pointer.
How you materialize your results for processing is your own business. The advantage of the relational model is its consistent throughout; the source data, the manipulation format for the operators, and the output are all relations. You can visualize it as "flat table" but that's an unimaginative visualization. You can just as easily twist that into a nested hierarchical structure in an object oriented language if you're so unfortunate as to be stuck working that way.
A "table" is only a visualization of the data, not the "form" the data actually takes and it's unfortunate that SQL chose this word.
It's better to conceive of a relation as a series of facts or propositions about the world. Each "row" is a statement. When read that way, it's a lot more elegant.
You can just as easily see nodes and edges in a property graph as propositions about the world. The nice thing is that you can model relationships between entities as first class entities. nodes have the implicit property of being non-fungible.
Do you know of any relational database that returns a query result as normalized tables the way neo4j returns a sub-graph?
Except network databases have little in common with graph databases...they're much more closely related to hierarchical databases.
I would say that Graph databases are now a strict superset of relational databases, not the other way around. In a graph database a node's named edge can naturally point to a node of any type or having any property schema. Doing this in a relational model requires one of several approaches that could only be classified as fighting against the model (or torturing it, as my PI liked to say).
The MIT team which is also part of the team that proposed GraphQL already kind of solved this problem using D4M [1].
Essentially D4M is able to universally represent relational, graph, spreadsheet and matrices using the mathematic technique of associative algebra, they even wrote an entire book on the subject [2].
It's beyond me why they didn't push forward for D4M but instead propose a limited GraphQL. It seems that they're restricting the D4M capability and focusing on query language like GraphQL, perhaps due to the industry demand and bias.
Heck the same team also proposed TabulaROSA, a new database OS as the more efficient alternative for conventional file-system based OS like Unix/Linux [3].
[1] D4M: Dynamic Distributed Dimensional Data Model:
[2] Mathematics of Big Data: Spreadsheets, Databases, Matrices, and Graphs:
https://mitpress.mit.edu/9780262038393/mathematics-of-big-da...
[3] TabulaROSA: tabular operating system architecture for massively parallel heterogeneous compute engines:
https://www.ll.mit.edu/r-d/publications/tabularosa-tabular-o...