Sysbench Benchmark for MongoDB – Performance Update
tokutek.com
tokutek.com
Perf Blog: http://www.tokutek.com/2013/03/sysbench-benchmark-for-mongod...
Download Source: https://github.com/Tokutek/mongo
TokuTek say to contact them for using it.
TokuTek made TokuDB, a storage engine for MySQL that uses Fractal Indexes (as opposed to B-Tree), read more at http://www.tokutek.com/products/tokudb-for-mysql/
What about read-heavy workloads, etc? Can someone with a bit more insight provide an objective overview?
The statement is really pretty much true. The best B-tree indexes (InnoDB) beat Fractal Tree indexes a little bit on in-memory reads, but we're not done tuning our implementation. On out-of-memory reads though, we usually beat B-trees because we compress so much better that we simply need to read less off disk.
Benchmarking is a tricky business though. Workloads can be varied, and you can probably find some corner cases where InnoDB beats TokuDB. I think we have some outstanding issues where the MySQL optimizer doesn't plan our queries properly, and we're still working on that. I'm a little out of touch with the MySQL side these days though so that could be wrong. But algorithmically, there are no cases where a B-tree index has a significant advantage over a Fractal Tree index.
Generally though, the read performance advantage we see is that if your indexes are Fractal Tree indexes, you can afford to maintain a richer set of indexes than you could have on InnoDB or MongoDB, and these extra indexes make your queries orders of magnitude faster. I think this is the most important (non-obvious) point to understand. I gave a talk about it here: http://www.youtube.com/watch?v=q6BnG74FZMQ
Of course what's not described in the paper include: transactions, mvcc, logging/recovery, parallelism, etc. Those are really interesting topics as well.
Regarding benchmarking, my recommendation would be to present other workloads as well; if what you say is true, you have nothing to fear, and rough parity with B-trees in your worst case scenarios would only make your argument stronger (it would definitely impress me more if it was included).
My point is that developers who are worried about index algorithm performance don't normally believe in free lunches, and your current benchmarks seem to be slightly cherry picked. Some of the wording is also a bit of a hard sell (and it sounds like your algorithm is good enough to stand on its own legs, so it's unnecessary). This will put off a lot of the more skeptical developers immediately.
Tim, our VP of Engineering, does most of our official benchmarking, in between project management, support, and testing (and a million other things). Just today (after we posted the TokuMX vs. MongoDB iibench blog), he said something about how he can't do many more of them, it just takes too long for MongoDB to finish the whole benchmark. That's time our servers can't spend testing the new software we're trying to ship.
My hope is that now that the code is open, more users will start pounding on it and publishing benchmarks. Justin Swanhart recently posted some great benchmark results where he showed us in a less-than-flattering light: http://shardquery.com/2013/05/25/tokudb-vs-percona-xtradb-us.... I haven't studied the workload he ran, but it looked like he may be pointing out some bad query planning behavior and slow cache warmup properties that we should probably address at some point. And that's good for the product and good for the ecosystem, so I am thrilled* that he posted it. I just hope we get 50 more just like it.
Anyway, I'm glad you've got a better understanding. Hope you can find a use case for it!
http://cdn.oreillystatic.com/en/assets/1/event/36/How%20Toku...
One cannot mix a replica set with TokuMX and MongoDB.
We have a very large install and migrating the data on production systems would be very difficult otherwise.
never mind - found some good answers here - http://postgresql.1045698.n5.nabble.com/Fractal-tree-indexin...
In all seriousness, this looks super cool, but I'm easily fooled by benchmarks and I don't have any projects right now that really push MongoDB that hard, so it doesn't impact me other than giving us a potentially faster Mongo, which is cool.