MariaDB acquires Clustrix
techcrunch.com
techcrunch.com
In 2010, they raised $12M in their Series B at ~$100M post-money valuation. Things were looking alright.
In 2013, they raised $16.5M in their Series C, and then shortly thereafter $10M in an unusual series D. That funding round reverse-split the outstanding stock 26-to-1 and converted all existing shares, preferred or otherwise, to common stock. What was left was $10M in new preferred stock, and $20M in existing common stock! New post-money valuation: $30M. This down round ended up being a 30x dilution for existing shareholders. If you had a tenth of a percent of $100M before, now you had a hundredth of a percent of $30M. Yowza!
After that bath, the board amended the charter so they stopped mailing out these notices. I don't know what's happened since, but I'll find out soon enough.
I feel bad that the company wasn't successful. It really was a great team and an impressive technical feat.
There has been a lot of work expanding the storage engine API so columnstore can be mainlined and work so spider can progress. I imagine that progress has showed them that they are likely to be able to integrate a much larger and less connected (to existing mariadb) code base like clusterix. I can only open they decide to open source (BSL license?) the clusterix solution eventually.
Between MyRocks (replacing tokudb), columnstore, spider and eventually clustrix seems that mariadb is trying to making the case that they can handle any size workload being through at it.
I tried to use MyRocks (never used it before) in MariaDB some months ago but couldn't find almost any docs and ultimately didn't understand which parameters were supposed to be set how under which condition... .
InnoDB isn't perfect, but it _is_ exhaustively documented and pretty well-understood, with a great set of related tools from Percona, etc, for simplifying operations. That goes a long way.
Recently we've switched back to using InnoDB for ingestion on one of our write-heavy tables and aggressively archiving the data out of it and into Clickhouse (InnoDB deals with the high volume of concurrent inserts, data is loaded into Clickhouse in large batches for querying). By comparison to Toku or RocksDB, Clickhouse is refreshingly well-documented and its easy for us to make consistent backups with ZFS snapshots.
I agree that tuning is too complex and we should do much better there. This explains where to ask for advice - http://smalldatum.blogspot.com/2018/05/where-to-ask-question...
It'll be interesting to see if this acquisition results in the interesting clustrix bits becoming libre software.
In-memory support was added later, though I'm not sure if it made it into a release as I stopped paying attention around when they started developing that feature.
They originally ran exclusively on specialized appliances having battery-backed write caches. The implementation is all in C using Linux AIO w/epoll. It's a sophisticated, high-performance storage engine.
When everyone shifted to the cloud and stopped owning hardware, Clustrix was forced to add a file-based backend. This still uses the high-performance AIO+epoll engine.
Single node failures cause entire cluster shutdowns, the cluster then takes forever to recover and must be done in a specific order. In fact just thinking about it makes me anxious.
Tables (and indexes) were automatically partitioned and replicated as needed, completely under the covers.
Queries (reads and writes) were distributed to the nodes where the data resided, in parallel.
Scaling the system was as simple as adding new nodes. Data was automatically rebalanced to take advantage of the new capacity.
Failure recovery was automatic too. If a disk or node failed, the data involved would be reconstructed from replicas and moved elsewhere with no interruption in service and no failed transactions.
It was a pretty impressive system, which predated Google Spanner. But, in the early days, you had to run their custom hardware to get it. There was no cloud version.
The video is from almost 5 years ago, but the high level idea discussed is still true today.