Retrospection and Learnings from Dgraph Labs
manishrjain.com
manishrjain.com
Some observations as a team active here.
We're adjacent / complementary because we make graph db's (and regular SQL/databricks/etc.) more actionable via rich & scaled graph viz workflows (analyst-facing) + graph automl (automation-facing), so sometimes end up working alongside graph db co's at enterprise/gov/tech customers so not just a silo'd db. I think we only saw ~1 production user of dgraph though, so entirely based on their competitors:
* Graph DB TAM: Neo4j revenue is probably around $200M/yr now, and probably another $100-200M across the other graph db vendors. (Market analysts say the graph db market today is $1B+ but seems rosy.) Importantly, because of analyst + AI use case growth in areas like recommendations, anti-fraud, cyber, etc., and streamlining of infra via cloud/docker/etc, everyone good is growing quite well, and I expect good YoY growth for everyone for at least 3 more years. The AI market likely a much bigger leap for graph, tho less clear for these graph db and especially cpu ones.
* GraphQL TAM: Agreed. But may be a bigger culture shock for a pivot. Not ready for Series B levels of expections. leading to...
* Revenue: Super risky, lacking big & growing revenue, to assume a Series A, and then a Series B (!), unless you have a special trick like having ties to the chinese government or being a successful serial founder people just trust.
* ... Tip for people looking at jobs: ask revenue / spending ratio + how many years in the bank when not profitable. If a b2b team can't make revenue work after $xM raised, they're on the path for stressful grinding, dilutive bridges & shutdown. VC treadmill grows expectations, so coming in from behind is asking for PE to take over or an acquihire where only the founders win.
I would agree that the "database" part is not that important. It is the notion of scalable graph analysis that sells a platform. However, for sufficiently large graphs (most of the interesting apps are in this class), GPUs often don't help much versus CPUs in my experience.
likewise, for analytics, we're seeing a bit of a split:
- DB side: Traditional real-time graph analytics (pattern search, ...) can often be done on extracts from KV-store level queries, so just use those and post-process the compute ("extract 1-2 hop neighborhood, then cross-product in lang xyz"). Like the old Titan -> Cassandra days. GPU can be nice for accelerating a simple inferencer here (e.g., a T4), but not necessarily the actual fetch. Gets a bit more blended in 'real' knowledge graphs, but that's more R&D. (Edit: I believe Facebook's at-scale graph engine successfully runs on top of SQL for related reasons.)
- AI side: Massive interest in GNNs (we're active here!) as starting to eat the lunch of the result quality that traditional graph analytics can give. Basically the pendulum swinging back to graph for areas like recommendors & classifiers. Shopping carts, fraud, cyber, etc. These had gone the way of ML + AI systems for awhile now, but with GNNs becoming practical, best of both. And... GPUs matter over CPUs again.
Funny enough, we're getting into a bunch of vector search scenarios, and because of our particular scale & query richness needs... looking at OSS graph DBs and pairing with GPU nodes. /insert "why not both" meme here
Which ones?
I don't understand what this could mean? How can a single hire be so impactful and so quickly? I guess the team was small but even then I'd be very interested to know how that could happen!
The two didn't get along and Manish was largely sidelined.
And then depending on who you talk to the company failed because of Gary (poor fundraising, strategy etc) or Manish (poor product, management).
The investors Dgraph had e.g. Redpoint, Airtree, Grok have an excellent reputation so there is definitely more going on then has been let on.
The worst thing about not being in the US is that there isn't enough public data points about which VCs are good or not.
And founders like Manish aren't stupid/drunk enough to ever spill the beans.
can you share some sources?
I can confirm the query language mistake. The Graphql-alike was different enough to be unfamiliar, and still felt somewhat awkward to use. The other query languagewas pretty badly documented and also wasn't great. Cypher isn't perfect, but it's pretty decent.
I also ran into a trivial bug around a comparison in a query not returning correct results and pretty much immediately gave up. (It don't remember the specifics, it might as well have been API misuse).
But the biggest issue with all these graph databases for traditional application development is schema. Most of them are schemaless or have some half-assed, basic schema support. Nebula is the only one I can think of with proper schemas.
This is not what you want for your primary data store! Even more so if logic is driven by JavaScript, or even Typescript since there are still plenty opportunities to mess up.
I wish there was a proper multi-modal DB that combines the best of both the relationalnand the property -graph models. The recently announced SurrealDb might fit the bill.
Graph stores are plenty popular for secondary workloads with special requirements, which is not dissimilar to Elastic search for the search domain.
I don't see that changing anytime soon.
So still in business although he seems lacking in senior management experience.
A few years ago, I spent weeks evaluating their products, including the graph db and some of their open source libraries. I have to say that you are a brave man if you use some of their stuff in your production system. Good luck to you and your whole team is the only thing I can say.
To give you some examples, when I had obvious safety issues raised, quite often I got pushed back with all sorts of excuses, including from Manish himself. That is something more than technical - it is a culture issue from the top.
From memory, they even had an angry ex-employee exposing all sorts of bugs, including safety ones, in their products after leaving the company. At some point, they even had their issue section of one of their github repos closed as a response!
If you don't know what I am talking about, just check their code committed back between 2017-2019.
FWIW I am really excited by surrealdb. I think it is the sweet spot that dGraph should have been.
I wonder if this sentence has similar basis:
From the blog post: "Dgraph took a hit suddenly due to a critically wrong hire — which made us go from a “things are looking great” to “sorry, you’re out” within a week."
I think this is sort of obvious. Imagine fixing a bug in your own personal project, compare that to fixing a bug in someone else's project. You can probably see the error and guess what the bug is, accurately, if it's your project. With someone else's code you're going to have to reverse engineer the system first.
Sustainable, open, software development feels like a problem that's still not solved. How do cloud offerings, consulting, extra features ("open core") and donations compare in terms of keeping development move forward? Have there been studies around that?
One model that seems to be workin is when companies have shared infrastructure and they collaborate on it (e.g. Linux).
Another one is when they use it as a standardization layer (e.g. Kubernetes and also note the much-different-nobody-remembers-about openstak)
Then you have open-core where you have a single entity (and the drama that comes with AWS picking it up) like ElasticSearch, Kafka, etc.
The other open-source + consulting (the RedHat model) has varied results because whoever consults gets incentivised to work against community in order to make money and the consultants will be in conflict too (Cloudera/Hortonworks, Mesosphere).
Not an exhaustive list - and not an expert.
It's likely worth mapping these "patterns". For any research I'd start from incentives with a game-theory causal approach.
Actually I hate the guts of cloud services. I hate adding another API key, account to my 1password, credit card, etc. I just want a Helm chart and be done with it. If anything, Kubernetes & GitOps is a way more ergonomic way for the vendor to get into your stack than stupid accounts are.
I was unfamiliar w this French phrase; here's a definition:
obligatory scene : a plot element that is standard for a particular genre
The biggest pain point for me was the query language schizophrenia: Incomplete support for GraphQL, or their custom Dgraph Query Language (DQL). As Manish said, they missed the GraphQL train. They really should have gone all in on GraphQL, and only GraphQL.
There may have been a slow-&-steady path to get to a mature state. But given the expectations of tech & vc scene these days, I can't blame the path you took. Good luck with the next venture.
I am glad that DBs like Postgres/MySQL, Sqllite etc got their freedom to evolve and slowly mature before the madness of fast-growing companies caught up to them.
If i saw a graph db with that as its primary language i would probably assume that its targeting a very different segment of the market than most graph dbs, and doesnt support deep recursive queries, and probably move on without a second look. I wonder if this sort of attitude hurt dgraph.