TokuMX 1.0: MongoDB with transactions, compression, MVCC, clustering indexes
tokutek.com
tokutek.com
Since it doesn't support cross-shard atomic operations, I'm assuming it takes the same approach as RethinkDB, i.e. co-locating data with its index entries, which means index lookups require querying every shard in the cluster. But I don't see this explicitly stated in the user's guide, so it would be nice to get some confirmation.
For 1.0 we wanted to match mongodb feature-wise as broadly as possible.
However, fractal trees can handle multiple clustered indexes very well, and this extends to a few ideas I have for adding things to do sharding in what I think will be a much better way. It's a bit up in the air right now but I can tell you we plan to make some significant innovations here, and if you have experience you can share, you should email us and we can talk it over and maybe change what we end up doing.
These are only initial impressions, but:
* It's definitely faster. Our write behavior is largely upserts against integers and doubles, and I'm seeing roughly 100% improvement against stock Mongo 2.4. The machine in question is an m2.2xlarge with a 1000 piops EBS volume attached, and it's doing about 7000 update operations a second. ( safe mode )
* I'm seeing consistently lower IO util than stock Mongo. Stock tends to vary wildly between 200-750 write ops, while under sustained write traffic, I see about 250 write ops.
* CPU usage is pretty well balanced against all cores, as opposed to stock's behavior.
* It's too early to say whether or not the storage savings will be as good as claimed, but at this point it seems that the TokuMX reprs are about 40% of the stock reprs. Like I said, most of our data is ints and doubles tho.
VERY LARGE CAVEAT I'm not running the database in a replica set because I'm lazy. So, the write throughput numbers are likely the best case scenario of what you'd actually be running in production.
(edit: line spacing)
If you've got to use MongoDB, it seems pretty nice. If it came with a tiny person that maintained the database for you as well, it'd be a no brainer.
Have you tried ObjectRocket (if you're in US-East or US-West)? http://objectrocket.com/ High performance Mongo with replica sets.
1. Is this a drop in replacement if I'm on Mongo 2.4? 2. If I want to migrate back to Mongodb from TokuMX, what needs to be done? 3. How quickly does TokuMX integrate improvements from MongoDB? 4. Is anyone using this in their production deployment?
How are you (Tokutek) planning to keep up to date with the MongoDB tree?
Are you planning on talking to the MongoDB folks about upstreaming this? Or will it be a pure fork with no sharing either way?
I have a few thoughts on this:
a) There is an astronomical difference between a neat technology and a complete, usable, supported product (especially with really complex software like databases). I can't tell yet how committed Toku folks are to this project. Is this a research project that may or may not go somewhere, or are they all in on the product? I think it's very important (for the customers and the industry) to get a clarification on this point.
b) I love seeing engineering projects like these. Experimentation like this (using a superb storage engine to power a popular db) is really exciting. I'd love to see where this goes.
c) RethinkDB has its own state of the art storage engine (with a very different architecture from Toku) that's tightly integrated into the full system. That lets us do very interesting things (fast path code paths, btree-aware caching system, etc.) The advantages and disadvantages of pluggable storage engines are really interesting.
d) If TokuMX does turn into a complete product, it's really exciting. It's nice to see the industry maturing.
We consider this to be full featured in the sense that we have a feature set that we feel users can deploy in production. As with any product, as users give feedback on what more they would like to see, be it existing MongoDB features or something else, we will use that feedback to enhance the product.
I completely agree that pluggable storage engines are an interesting topic. But we went the integrated route (ie: no storage API) probably for the same reason: things get simpler and easier to implement when the stack is shorter.
As for licensing, TokuMX embraces the spirit of open-source and we're confident our open-source licence plays well with the AGPL.
Originally it was used with ISAM/MyISAM and it was pretty popular. Then InnoDB came around and it quickly revolutionized the MySQL world, allowing MySQL to grow to the next level. Now InnoDB is by far the most commonly used storage engine and the default on several distributions.