Cayley – An open-source graph database
cayley.io
cayley.io
I'm curious because I've had a couple situations where I thought using neo4j (or some graph db) would be a natural fit for something I wanted to do, but otherwise I thought most of my other data fit into postgresql just fine. My instinct is that if I'm doing this in a web app then querying from two different databases is going to slow down my responses a lot.
I use both, in a similar vein of "using Elastic Search". It could be your primary store, but it's sometimes it's more pragmatic to have two sets, a "solid" base.
This is not to say that it can't be done. What I'm stressing is that larger "changes" are hard and difficult to handle - which means a lot in the start of your process, and less in the end, as in, when you're deciding how to model your data. For instance node layout (new properties? different type? other constraints?), and mass updates are also a bit cumbersome.
Usually I have more than one SQL table (naturally) since the data I've used in graph databases is mix and match (otherwise I'd just use a fixed schema and some relational DB).
-- As for "how that works", for me it's:
routinely update from my base database with queries alike: ID > last ID.
This has worked as expected, in terms of what data you get in, and which limitations you impose (e.g. timeliness).
I'm currently making a shift to running all data in my graph database as I've settled on a model (which edges, which nodes, which properties).
> querying from two different databases is going to slow down my responses
True, but depending on your data (do you know one of the queries beforehand - e.g. is your postgresql query enriching whatever your graph query returns) you might have success tying (inserting) some of the SQL data to your graph database.
It led to itself as neo4j was only useful for parts of our queries and fitting everything into neo4j was just hassle when most of our data was relational.
> querying from two different databases is going to slow down my responses
I think querying 2 different systems tends to be slower, but more importantly you lose transactionality. If you can use a single system that is at least on-par with your relational system for your run of the mill data and have a very powerful graph then that's a big win.
We use the PACER engine (https://github.com/pangloss/pacer ) to power queries.
This approach allows you to get the optimal performance and only reaching out to other systems when needed.
Some projects I use NoSQL, mongo.
I do use Redis for reactivity and caching (but redis isn´t a database so)
Congratulations for the progress, thanks for sharing your work. I do want to see more great odbs with friendly APIs
1. Adding edges to a neo4j graph is a painfully slow process. For a large graph with a few million nodes - it'll take days. 2. Scaling neo4j on a cluster is either not possible or it's a painful process. I'm yet to discover this.
However, the greatest advantage that neo4j offers is the ability to query a path. So far, no other graph databases that I know have this ability (including Apache spark and giraph).
It's quite possible to build a directed graph database as an adjaceny list in redis. We tried this and it's super fast and scalable. However, querying is very painful.
2. The reason Neo4j is the only database that allows you to query a path is the same reason that setting up clustering or sharding is difficult. If your graph is complex then the problem is "How do I split up these subgraphs into shards so that traversals don't have to traverse across shards?" -- Building a giant adjacency list and using that as a traversal index is a clever idea, I must admit. :)
Virtuoso's support for open standards makes it easy to use it as a complete solution covering all the bases, or, as in @philjohn's case, to plug-and-play with best-in-breed solutions along any axis where our implementation proves not to serve your needs for any reason. (We do want to know how and why we don't measure up, so we can improve that aspect!)
Graph stores are typically much slower for repetitive data that fits cleanly into a relational model. This isn't to say they're not useful - for more irregular data they're a fantastic fit - it's just that very irregularly structured data isn't the common case.
Of course, you can always use two different stores - much like many sites do with a separate lucene/elasticsearch index for text search - but your graphing needs must be relatively componentised for that to work well.
So an ORM that usefully interpreted model subclassing etc. and created self-joining tables and could query the resulting model using RECURSIVELY WITH as appropriate would be a real boon.
I was looking into using Stardog for a metadata repository I was building, but we ended up (probably unwisely) bastardizing Postgres into a bunch of self-join heirarchies.
My experience is that if you don't need constraints/enforced relational integrity, RDF stores make for really simple/easy object storage. There's definitely a performance tradeoff, though - depends on what you need, really!
// Our triple struct, used throughout.
type Triple struct {
Sub string `json:"subject"`
Pred string `json:"predicate"`
Obj string `json:"object"`
Provenance string `json:"provenance,omitempty"`
}(actually I've just made that up)
if it is for production, how is your read/write performance?
And on the topic of Cayley... I'm stoked. I'll give it a try tonight. Easy to use, and graph databases, haven't really gone hand in hand for me.
Did you want to get support from the company that makes neo4j? If not, then you don't need a license.
Where using more than one instance would be desirable (I'm not trying to over optimize-- I swear! ;-_-). I suppose one server can go pretty f#€£ing far. But all the clustering requires a license (~$12k per server for start ups?) unless the code which uses it is GPL or AGPL (I think?).
Anyone who wants to explain the AGPL in layman terms would be greatly appreciated. Or specially Neo4j's application of it. To me, it seems cost prohibitive for lean startups, of one or two people.
Again, they've said contact 'em-- so it may be case by case. I could just be worrying about non-existent problems.
I tried Titan and Orient, I didn't find either as nice as Neo4j to use in my code. Though Neo4j, for me, was much more difficult to setup and use as a cluster (at the time).
... And more on topic, I'm about to install Cayley!
(Summary of the Affero GPL: you must distribute the full source of your service, under an AGPL-compatible FOSS license, to the users of your service. Summary of the GPL: you must distribute the full source of your program, under a GPL-compatible license, to anyone you distribute your program to.)
It does look like the "Community" edition leaves out the clustering features, so if you need those you'd likely need a license for the proprietary version.
It's pretty much useless in a server environment since even replication isn't possible; if you want a redundant setup -- which you will -- you will have to keep the nodes synchronized yourself. Not to mention that since Neo4j is an in-memory database, it puts a hard limit on your dataset size.
They have a "Personal License" [2], but it lasts for one year and you're not allowed to use it if you have capital funding or a certain amount of revenue.
I was actually referring to the "High-Performance Cache" mentioned in the feature matrix.
> Neo4j Enterprise is available for free for any AGPL project
I was really talking about a commercial setting. How many companies deploy a fully open-source project (ie., honouring the requirements of the AGPL) in a redundant data center? Not a lot, I imagine.
> Neo4j is not an in-memory database
True, it seems I was misinformed about that.
Cypher's not bad either, but it's not "just Javascript". But I'm totally taking commits if someone wants to port it :)
http://google-opensource.blogspot.hk/2014/06/cayley-graphs-i...
Some of the things I like about cayley
1. Switchable backends(I wonder if I can configure it to use couchdb as a store)
2. Documentation get right to the point. When I first tried my hand at graph databases I could not understand where to start but cayley's approach is pretty straight forward and it wins plus points from me for including a big dataset :)
A question: I see no mention in the docs about running it on multiple nodes (where does it stand with regards to CAP etc)
Interesting to see more products entering this space.
I played with it, and it was kinda fun.
Deleted comment
However, I'll point out that most of the actually-used open-source projects coming out of Google aren't actually "Google projects", they are "projects by Googlers released under the OSPO process". LevelDB (Jeff Dean & Sanjay Ghemawat), Protocol Buffers (Kenton Varda), Guice (Bob Lee, Jesse Wilson, and Kevin Bourillion), Gumbo (myself), and angular.js (a team within DoubleClick) all started out as small internal projects built to scratch an individual's itch that were then released externally because hey, why not. The "corporate" open-source projects have been things like Android, Chrome, GWT, Closure, Polymer, etc.
Does it have bulk import and if so what is it's speed for bulk import rougly speaking?
Load speed is pretty good (into persistent storage, I assume) and can be improved with some of the database parameters. A rough estimate is that a million triples or so takes about 5 minutes, but that slows down as it gets bigger. 134m triples took me 6-8hrs, so I slept on it.
It _does_ support bulk loading, with over 500K triples per second. According to http://franz.com/agraph/allegrograph/agraph_benchmarks.lhtml, given enough RAM, it can load over a billion tuples in just over half an hour.
Not a Google project, but created and maintained by a
Googler, with permission from and assignment to Google,
under the Apache License, version 2.0.- Jonathan
OS X: Homebrew is the preferred method.
Over my dead body. Homebrew messes up my /usr/local, leaving its heaps of crap everywhere. I don't find that acceptable. MacPorts at least has the decency to put itself into /opt/local, which isn't normally used by anything else.