Introducing OpenCypher, the open graph query language project
neo4j.com
neo4j.com
> I believe that graph query language is Cypher.
I, and others, believe gremlin is that common graph query language.
My first meet with graph databases was in fact Neo4j, but quickly pivoted to Titan, and now using Cayley for smaller projects (although its Cassandra model will make an interesting future).
All except Neo4j supported gremlin, which to me is expressive, formal and actual-human readable.
Cypher looks cool, and has intuitive method of writing the edge query parts, but while it looks hip, I didn't find it obvious nor human parsable. Human readable yes. You recognize the terms, but translating it into an AST is simpler and easier in gremlin than in Cypher (in my opinion)
Gremlin was originally developed by Neo4j co-founder Peter Neubaurer and Marko Rodriguez (Titan founder) while they were both working at Neo4j.
SELECT ?cypher_attributes
WHERE {
?cypher a <QueryLanguage> ;
<queries> ?graphs ;
<attributes> ?cypher_attributes .
?user <USES> ?cypher .
FILTER (?user IN ‘Oracle’, ‘Apache Spark’, ‘Tableau’, ‘Structr’)
?opencypher <MAKES_AVAILBLE> ?cypher .
}
Instead of MATCH (cypher:QueryLanguage)-[:QUERIES]->(graphs)
MATCH (cypher)<-[:USES]-(u:User) WHERE u.name IN [‘Oracle’, ‘Apache Spark’, ‘Tableau’, ‘Structr’]
MATCH (openCypher)-[:MAKES_AVAILBLE]->(cypher)
RETURN cypher.attributes
In the SPARQL case the graph flow does not revert on the edge with <USES> (it can using
?cypher ^<USES> ?user, but it would be weird and in the larger queries very confusing). The SPARQL case also tends to group related concepts together.This assumes a DEFAULT BASE URI is selected for the SPARQL version that contains all the modeled relations. Which in a straight comparison to Cypher is a fair comparison.
I find Gremlin a lot nicer than Cypher, and a lot more powerful as well. Also up to today Neo4J just has not scaled all that well. I am awaiting the LDBC Benchmark results of Neo4J to see if I am wrong.
What Neo4J has been great at is making a nice solid product that aims at solving developer problems. I believe as a database it has not been that great at solving enterprise or life science community problems. Its still a single database instance without federation on demand.
MATCH (openCypher)-[:MAKES_AVAILBLE]->(cypher:QueryLanguage)-[:QUERIES]->(graphs),
(u:User)-[:USES]->(cypher)
WHERE u.name IN [‘Oracle’, ‘Apache Spark’, ‘Tableau’, ‘Structr’]
RETURN cypher.attributes
I guess, in this particular case, it's subjective preference which language you feel expresses the query pattern most legibly. I certainly prefer the visual approach of cypher.In Gremlin3, the above is:
g.V().match(
as("openCypher").out("makesAvailable").hasLabel("QueryLanguage").as("cypher").out("queries").as("graphs"),
as("user").out("uses").as("cypher"),
where("user", within(["Oracle","Apache Spark", "Tableau", "Structr"])),
select("cypher").by("attributes")That can of course also be affected by my slight reading disability where the shape of the words is important. This shape could be disturbed by the connecting sigils in both gremlin and Cypher. So I understand that my preference might not hold for the whole population :)
g.V().match(
as("cypher").hasLabel("QueryLanguage").out("queries").as("graphs"),
as("user").out("uses").as("cypher"),
where("user", within(["Oracle","Apache Spark", "Tableau", "Structr"])),
as("openCypher").out('makesAvailable").as("cypher")).
select("cypher").by("attributes")
The original query is sort of an odd query as you don't don't need all the unbound variables...And if you want to use SPARQL over TinkerPop, just use a SPARQL->Gremlin Virtual Machine compiler. https://github.com/dkuppitz/sparql-gremlin
Do you have an example of a query that you think is better in Gremlin? I've yet to see one but haven't spent much time with Gremlin.
http://speakerdeck.com/timwilliate/graphs-are-feeding-the-wo...
Many meaningful lineages in life sciences can be hundreds to thousands of levels deep (our datasets are great examples). Neo4j is the only graph database I have evaluated that handles traversals across lineages of this depth while still achieving the performance scalability promised by maintaining index-free adjacency across which ever node in the cluster a traversal is sent to.
I am not saying that Neo4J is a bad choice, I am just saying that it due to its lack of federation support it is an expensive choice for the life sciences. i.e. an economic argument over a technical one, and not even looking at 1 project a time but in general for the community. Neo4J and Cypher will never support federation in the way that SPARQL allows. This is because all this URI business in RDF is annoying when modelling your data but critical when merging datasets on demand between separate databases. e.g. joining ChEMBL & UniProt & MeSH & PubChem etc...
We in the life sciences rarely do graph traversals for graph traversal sake, but tend to join trees. e.g. intersect a branch of a taxonomic tree with a branch of the GO tree. There are cases where real graph traversals are being done (assembly&variation graphs).
OpenCypher is a great step forward. Now Neo4J needs a open public standard for serializing graphs to disk that can imported into Neo4J and other databases. RDF being supported by so many different databases allows us to support many more of our users (at UniProt) even if they don't use SPARQL or our choice of Graph database themselves.
I'm curious to see what would be the popular vote of what query language the industry wants to standardize (I definitely hope not GraphQL, I'm personally think RQL and LINQ are interesting). Rather than some company attempting to standardize it, I want to see the users vote. What would you choose?
From https://en.wikipedia.org/wiki/Cypher_Query_Language:
MATCH (charlie:Person { name:'Charlie Sheen' })-[:ACTED_IN]-(movie:Movie)
RETURN movie
to SELECT movie
FROM Movie movie, Person person
WHERE person.name = 'Charlie Sheen'
MATCH person-[:ACTED_IN]->movie;For that pattern to work in SQL, your SQL engine would need to be able to view its relations (tables) as a graph. I'm not smart enough to sort this out in my head, but my instinct tells me that it's better to have the two models separate, mixing them in one query seems like a recipe for confusion.
g.V().has("name","Charlie Sheen").
as("person").out("actedIn").as("movie").
select("movie")
However, this can be expressed in a much simpler form as you don't need all the variables. Simply do: g.V().has("name","Charlie Sheen").out("actedIn")Using JSON for the template also leads to a clear correspondence between the query and the result which is a very nice property few query languages have.
http://mql.freebaseapps.com/ch03.html
to take trip back to 2007.
tsturge did a prescient job with the original (yeah, GraphQL is 2007 all over again) and I wanted to keep the flame alive.
Back then both were pretty rough, Gremlin looked more polished but wasn't really (most promised features in the documentation were not working just yet).
What has changed since?
http://tinkerpop.incubator.apache.org/
The two big things:
1. Gremlin language is much cleaner and easier to use.
2. It supports OLTP graph databases (e.g. Titan/Neo4j/Stardog) and OLAP graph processors (e.g. Hadoop/Spark/Giraph).
Here are some use cases I think Cypher expresses nicely that I (as a GraphQL noob) don't know how to do in GraphQL:
Simple recommendation engine - suggest people with lots of friends in common that I don't already know;
MATCH (me:User)-[:KNOWS]->(friend)-[:KNOWS]->(fof)
WHERE NOT (me)-[:KNOWS]->(fof)
AND id(me) = blah
RETURN fof.name, count(friend) AS friendsInCommon
ORDER BY friendsInCommon DESC
Basic routing - what's the shortest way for me to get to work? MATCH p = shortestPath( (home:Address)-[:ROAD*]-(work) )
WHERE home.street = .. AND work.street = ..
RETURN p g.V(blah).out("knows").aggregate("friends").
out("knows").where(not(within("friends"))).
select().
by("name").
by(count("friends")).
order().by(valueDecr) MATCH (me:User)-[:KNOWS]->(friend)-[:KNOWS]->(fof)
WHERE NOT (me)-[:KNOWS]->(fof)
This easily reads as "get friends (AS FOF) [who know] friends [who know] me, where me [does not know] FOF" to me. out("knows").aggregate("friends").
out("knows").where(not(within("friends")))
The Gremlin, on the other hand, reads as "get friends [who know] friends where... friend is not a friend???" to me.I also don't easily see where something should be a method and where it should be a function. Why not `order(by(valueDecr))`? Why not `select("name", count("friends"))`? Why not `where(not().within("friends"))`?
SELECT ?foaf
WHERE
{
?me a <USER> .
?me <KNOWS> ?friend . ?friend <KNOWS> ?foaf .
MINUS {:me <KNOWS> ?foaf }
}
OR SELECT ?foaf
WHERE
{
?me a <USER> .
?me <KNOWS>/<KNOWS> ?foaf .
MINUS {:me <KNOWS> ?foaf }
} out("knows").aggregate("friends").out("knows").where(not(within("friends")))
OR out("knows").aggregate("friends").
out("knows").where(not(within("friends")))
Note the "." concatenation that ties the two lines together into a chain. When nesting parallel traversals (e.g. match()), the traversal patterns are delineated by ",". . = AND
, = OR
Ha. Thats a generally neat way to think of "." and "," in computing. mult and + ...the algebra.It's better to understand GraphQL as a protocol that competes with REST. It's only a language in the sense that JSON is a language; i.e., it has a syntax that can be parsed.
For example, GraphQL supports queries like this:
query movie {
whereYear(max: 1985)
actors {
hasName(like: "goldblum")
}
}
But this is something the particular schema and implementation would need to implement. If you want to filter by arbitrary attributes, you're out of luck because the spec is just a syntax. I suppose you do something like: where(what: "year", max: 1985)
but you still have to invent a standard set of parameters here: min, max, eq, notEq, lessThan, lessThanOrEq, like, etc. Again, totally ad hoc.GraphQL, not being a language, also doesn't support variable bindings. So you cannot do self-referencing queries like "find all movies with a director who also acted in it", because that would require some kind of variable support.
(This is not a criticism of GraphQL, by the way. It's great at what it's defined for.)
So, for people more comfortable with SQL, your question could just as well be "what can SQL do that GraphQL can't"?
The answer is, of course, that they occupy different domains and have different functions. Both can do lots of things that the other can't.
I find what they have defined so far far more approachable than Cypher or Gremlin. As they have been adding features it is starting to sprawl and look just as nutty as the others. But I do like how it is defining the whole ecosystem around how graphs can be defined and interacted with, much like Gremlin has, but with a much more focused and disciplined approach.
GraphQL has nothing to do with graphs.
OpenCypher has nothing to do with ciphers.
Cypher is named for the character in the movies. :)
Between the orientdb/neo4j dick swinging contest and the whole oreilly "graph database" book for pure fluff. I am left with a very bad taste in my mouth.
It tries to do the whole vendor lock-in thing... badly.