2021 in Database Startups: Gold Rush
unum.cloud
unum.cloud
That drives up valuations in the short term and results in a very easy time to raise capital (aka Gold Rush). However what follows is a period where the market will turn the screws really hard on companies to show they have a viable and economically justifiable business or they’ll get washed out real quick.
We’re not in a bubble like dot com but it will get very uncomfortable for companies that don’t have a clear path to self-sustaining economics.
After that $CHWY managed to go as high as $118/share, translating into a market cap of roughly $50 Billion in December 2021. Now they are trading at about half that price.
[1]: https://www.cnbc.com/2019/07/26/opinion-chewy-is-no-petscom....
- JSON / jsonb
- improved partitioning support
- identity columns
I agree that I wouldn't want to trust a new entrant to the space unless it was subject to rigorous testing, eg https://jepsen.io/
Postgres has the advantage that it's still open-source, can be ran locally and self-hosted (and has a healthy ecosystem of managed providers) and won't go away anytime soon. None of these are true for all these new DB startups.
- The product has bugs not addressed by Jepsen.
- Performance is inadequate.
- The optimizer is immature, and you find yourself struggling to understand and work around its limitations.
- Replication is missing or doesn't work well enough.
- There is absolutely no substitute for N years of real-world experience, no matter how rigorous the company's testing.
- Impossibility of finding people already familiar with the product, because they are all employed by the vendor.
- Even aside from product issues, is the company profitable? If not, how long is it going to be in existence?
I've actually seen the results of a database product being picked, the company behind which ceased to exist. It wasn't pretty, since at one point the product ceased to be supported and therefore neither any updates/fixes were made, it wasn't available in the repositories for new OS distros and eventually even the documentation for it went offline. Having to support a system that integrated with it was an unpleasant experience, all the way to it being eventually replaced with something else instead.
Therefore, it probably makes a lot of sense to base something as critical as your data storage layer on proven technologies that have demonstrated that they'll probably be supported in one form or another for the following years or even decades, unless you have a good reason for choosing something else.
Those reasons might deal with particular workloads or requirements, e.g. clustering solutions for PostgreSQL/MySQL/MariaDB/..., geospatial extensions, solutions to integrate with it through REST interfaces or even GraphQL or something like that, with a stable and proven piece of software still at the core of it all.
In case anyone is wondering, the product in question was Clusterpoint, about which you can read a bit more here: https://en.wikipedia.org/wiki/Clusterpoint a NoSQL database that actually predates MongoDB by a few years, as far as i know. Of course, now it seems like even their homepage is offline.
> Clusterpoint Ltd.
> Founded 2006
> Clusterpoint Database
> Initial release 2006
> Stable release 2015
MongoDB: https://en.wikipedia.org/wiki/MongoDB
> MongoDB, Inc. (formerly 10gen, Inc.)
> Founded 2007
> MongoDB
> Initial release 2009
> Stable release 2021
Then again, maybe Clusterpoint's Wikipedia page is lying, i don't really care much about it anymore.
The second one would be keeping your promises or even competing entrenched databases.
Even when you have gotten this far a way that you can exit is to get one enterprise customer convinced to use it so badly they must acquirer you.
The biggest lie I have seen across databases is actual time travel for your data. Or in other words given a date and time in the past it gives you what the data was then.
I worked at a company where we built this for our internal HBase clusters. Given any timestamp (or even better a commit ID) we could restore the cluster to that point in time. This was done through a combination of backing up the store files + write ahead logs. As I recall we had a retention limit for how much time we could do this for - but they were high enough for all of the "oh crap" moments.
Did they integrate with an underlying filesystem's snapshotting feature to make the store file snapshots differentials, or was that also implemented in the database? Or just threw space at the problem?
I have a sketch of the lang side here:
A big problem is that get funding outside of USA is not that easy, and here in Colombia this kind of projects is not considered "profitable" by investors.
But still convinced exist a lot of untapped potential in this area!
Databases in 2021: A Year in Review - https://news.ycombinator.com/item?id=29731885 - Dec 2021 (126 comments)
The first is RethinkDB [1], which was a darling for awhile but couldn't seem to find a business model. I really wonder if it came later how much money it could've raised. There were claims made that the cloud killed RethinkDB. This past year seems to fly in the face of that theory.
The second is: is this just database "success" (as measured by funding rounds and valuations) or is this just the general case with all startups? I honestly don't know.
Still, I wasn't aware of all these players and it was a good post so thanks for that.
For DBMS startups, which don't offer an order of magnitude improvements, the problem is convincing any one to replace a fundamental piece of infrastructure. For disruptive startups the problem is even bigger. As soon as you go open-source, all of your competitive technological advantage is gone.
They have a new PAYG serveless offering that starts at $0 and scales up.
I’m not sure what BLAS has to do with a key-value store.
Hahaha, thanks for the kind words! Not sure if I could describe our methodology better)) We originally come from and hope to continue working on AGI technologies (Artificial General Intelligence).
Writing a key-value store and, subsequently, a DBMS came out of necessity, as other solutions seemed too bulky, slow, outdated and expensive. So we invested time into building UnumDB for internal use, but after first benchmarks, it was clear, that this must be available to broader audience. We will polish it in the next couple of months and make it deployable in public clouds, just like other DBMS brands.
Now DBMS is our central focus, but with essential R&D out of the way, we will rebalance our priorities and continue working on BLAS and sparse neural nets.
It was on the longer list of smaller DBMS brands, but I was afraid to make the volume incomprehensible. Maybe we will write a more focused version, covering just smaller startups in 2022.
Arango is another one, seems to target Germany maybe kind of like SUSE