PathQuery, Google's Graph Query Language
arxiv.org
arxiv.org
BTW, I think this paper is very well written. Graph queries is not an easy topic and their examples and text make the data types, syntax, and how to perform matches and aggregations easy to follow.
I used to be a huge Prolog fan, but I don't think that anyone has paid me to do Prolog develop in over 20 years. My largest Prolog project was porting a prototype AI planning system that I wrote in Common Lisp to ExperProlog. It took about 6 weeks to prototype the system in Common Lisp, and re-writing it in Prolog only took about two weeks.
I noticed last month that someone was offering an online Prolog course using Swi-Prolog. That might be a good place to start.
Datalog solves most of that :) My read of the paper is their semantics largely reduces to datalog. The most interesting semantic exception is their bounded recursion, which suggests they're doing optimizations normal datalog wouldn't. Likewise, graph workloads often have weird phenomena to optimize for in terms of expected queries & data, where I'm guessing datalog might fit semantically but not how people would normally implement an engine. The PathQuery paper carefully side-stepped all such discussion, so I've been curious.
At Graphistry, we do end-to-end GPU stuff, so I've been thinking about this particular problem, and came to a similar conclusion of a datalog-ish subset being the most straightforward way to get more performance/$ for this kind of task.
https://opensource.googleblog.com/2021/04/logica-organizing-...
(Datalog/Prolog family language compiled to SQL)
And yes, this kind of thing is why datalog is a lot more amenable to fast query plans & runtimes than prolog. This part is especially cool: https://github.com/EvgSkv/logica/blob/main/compiler/dialects...
Warren started working on PathQuery as a replacement for MQL. PathQuery was originally executed on Pregel and was far from a realtime query language.
A realtime query engine was later developed for PathQuery in the search stack so interesting search queries could be answered a la minute from the Knowledge Graph, rather than just pre-generating results for a limited class of queries.
This was all done in the 2012-2013 timeframe. Most of the people involved are long gone.
I think what differentiates PathQuery from many "query languages" is that it takes on the task of transforming that data into somebody else's schema (aka "turning protos into other protos"). Having a DSL for this is particularly attractive in a company that will otherwise expect you to do string formatting in C++. But even under normal circumstances there are wins from having data transformation vertically integrated with your query language.
The semantic web feels like a fad that has clearly passed but there are plenty of good ideas from that field that are under-utilized. For example, a lot of what gets encoded as JSON can be described with these triples and once the data is in that form, there are interesting questions you can ask of it.
@entities
.[/type == (Id(/ museum ) , Id(/ theme_park ))]
.{
id: ?cur
require name: /name.[TextLang() == en ]
@merge : events::GetInfo()
}
can be written something like: *[type == "museum" || type == "theme_park"] {
id,
name: select(name[lang == "en"] => name),
...{
// GROQ doesn't have functions, so GetInfo()
// would need to be inlined here
}
}
PathQuery is of course a lot more complex, but the basic structure seems very similar. One thing GROQ does not have yet is recursive querying of the type needed to traverse graphs, but this is on our roadmap to implement.We've published a public draft specification for GROQ, and hoping it will be adopted by more tools. For example, we already have an open-source JavaScript implementation of it.
(Disclosure: I work on GROQ at Sanity.)
So the paper has a sort of weird feel. It goes like this: PathQuery is a good language, because it has good semantics and it is optimizable. Some words on its good semantics. We don't discuss how it is optimizable, but trust us, it is.
*[type == "user"] {
id, name,
"slug": id + "-" + lower(name),
"photos": photos[] {
url, width, height
} | order(position),
"newestComment": *[_type == "comment" && author == ^.id] | order(createdAt desc)[0] {
id, body
}
}
PathQuery seems to give you similar tools in transforming data as part of your pipeline, which I think is how query languages should be like.To support multi-locale names in Datomic or Datascript-flavoured Datalog attractions_v1.pq would be written as (where :entity/locale is of type :db.type/ref):
(d/q [:find ?id ?name ?lang
:in $ ?type
:where
[?e :entity/type ?type]
[?e :entity/locale ?loc]
[?loc :locale/name ?name]
[?loc :locale/lang ?lang]
db #{"museum" "theme_park"})
=> (["/z/38dwfnb8" "Museum of Modern Art" "en"]
["/z/38dwfnb8" "Museo de Arte Moderno" "es"]
...)
You could store the locale inside the name string, but then you can't efficiently filter on it. If you want the whole shebang, can do `(pull ?id [*])`. It's not obvious in their example if the locale filtering is post-query or part of the unification.With only 4-tuple EAVT-style indices, you need an intermediate join to support compound values like ["Hello" "en"] if you want to query against the components, so I have been toying with making my own graph DB with 5-tuple (or more) extensible indices, e.g. EAXVT, where X would be the locale, but then index order matters.
I may be missing something though.
I've never built a language, so maybe I'm just crazy, but what's wrong with, say,
https://gist.github.com/insanitybit/cd997e8d367889708edc3ec2...
Basically I defined the subject in a sort of 'namespace' of constraints ("parent", "children", etc), and then the edge or predicate constraint on it.
I guess it's easy to just "make up" a language, and implementing is the difficult part, and then it's like "ok well then how do you solve <some unanticipated problem because idk what I'm even talking about>" but also, graphql, as much as it isn't good for a database API, is also quite straightforward and familiar looking and I feel like taking something like rdf as a basis would be reasonable.
https://github.com/JeffreyBenjaminBrown/hode/blob/master/doc...
It still seems to be missing some things. For example, they show a recursive query, but no way to return the actual path taken, only the start and end points (I also can't figure out how to do in SPARQL); I don't see a way to explore predicate relationships; is there a method for reification?
On the other hand the sparql in the paper is decent but not excellent. The result is subtly different from the path query one, but I think fine for example use case.
The only query languages that do this "right" are datalog (Datomic) and Q/K (KDB+, Shakti).
After working on 2-3 iterations of similar languages (sexpressions, clojure enhanced with pipe operator), my take away is to define an API (not a new language) and support multiple languages on top of it.
I picked python and the API is here:
https://github.com/adsharma/fquery/ https://adsharma.github.io/fquery/
Like SQL, DDL and DML should be separated too.
But in and of itself, GraphQL isn't really built to query graph databases as you said.
[0] https://dgraph.io/docs/query-language/graphql-fundamentals/
Any graph, when you traverse it from a specific node, with fixed depth, and you don't explicitly work with node references, but rather node "values", looks like a tree that sprawls from that node.
Anyway, PathQuery is a completely different beast regardless.
A graph query language implies you're querying a graph with specific edges, not ones constructed on the fly from the query.
GraphQL's goal isn't to be "advanced" either. Its goal is to allow access to any data structure (be it backed by graph, RDBMS, file system, NoSQL, etc. etc.) without burdening said structure with capabilities that are not characteristic to it.
Like, if you query a graph, the idea the query may provide any arbitrary join expression would completely drive that graph's performance into the ground.
And not automatically joining by foreign key is a mistake that's going to haunt us for a while I bet.
However when we're arguing capability then it's just as capable of querying a graph as graphDB is. More if the version of SQL supports recursive joins.
And no, arbitrary join expressions are not a feature that "haunts" SQL databases, because unlike graph databases, relational databases ARE built for that, and it's one of the primary reasons SQL databases are very resilient to change in face of constantly changing ad-hoc query requirements. And it's an important feature of relational algebra that is used every day by countless applications.
SQL and GraphQL serve different purposes at different application layers. Both do precisely what they have to do. The fact they're a bit similar is not coincidental, but also they're not mutually replaceable.