RethinkDB 1.7: hot backup, atomic set/get, 10x insert performance improvement
rethinkdb.com
rethinkdb.com
https://github.com/rethinkdb/rethinkdb/issues/97
It's a little disappointing that it has not been resolved for several releases now.
Is there a target milestone for a production-ready release of RethinkDB? Is it 1.8 or 1.9?
RethinkDB will be marked production ready when it hits 2.0, though it will be ready for useful work way before that. (A huge number of people are already using it in their daily work and building important software on top of Rethink)
They even sent me a small care package with a hand-written note, it was great of them. Marc, the guy who sent me the package, is now one of my regular DOTA2 teammates, and he's great at that too.
Major props to them, I really hope Rethink becomes as amazing as I think it will. I'll give it another spin over the weekend and write it up.
Congrats on the new release, guys!
I'd encourage you to learn more about that movie.
1.0 is supposed to mean production-ready. Are you suggesting this should be a 0.7 release?
Version numbers don't have on single universally accepted set of defined semantics.
Not directly related to the new version, but speaking of upgrading previously, make sure you do the migration scripts beforehand if you have a current server.
I make the mistake of upgrading and not migrating beforehand from 1.4 to 1.6 and couldn't find an version of 1.4 since all the archives were down and building from source on a VPS just wasn't happening. The Rethink team was amazing in their support of building the old version for me specifically, and if this is indicative of their dedication to user support can't wait to see this become a big success.
This kind of dedication to it is definitely what makes me want to check it out, you guys are doing some really great work on it.
http://stackoverflow.com/questions/15151554/comparing-mongod...
MongoDB: 0m0.618s
RethinkDB: 2m2.502s
Although as stated in that post, this is likely because of the different fsync() policies between the two databases.I suspect that RethinkDB 1.7 would be comparable to Mongo on insert performance workload described in this stack overflow thread.
We're working on getting authoritative benchmarks out, but unfortunately good benchmarks are extremely time consuming (much like any science experiment).
/dev/null: 0m0.001s
/dev/null is clearly the best database.(If your performance numbers are too good to be true, they might not be true.)
There's also the angle that if it's to offer no performance benefit, perhaps a classic relational database will do it?
I'm curious to hear your thoughts about it.
According to iostat, this does 5000 write transactions per second to the disk.
If you are logging throw-away data, you can use an UNLOGGED table. This brings it to 49,000 inserted rows per second.
If you don't care about your data integrity, and set fsync=off, it goes up to 65,000 and seems to be limited by CPU.
It uses a single process/thread/connection. Adding more concurrent clients would improve performance, though there is still some scalability work to be done.
Ops/second in the graphs measures number of documents, not groups.
One fundamental aspect of the benchmark is that it uses random keys, which is significantly slower (in any database) than using sequential inserts. I'm not sure how you insert into Postgres wrt to keys.
I suspect postgres has better latency, and since the benchmark in the post is very much latency bound, this would account for the factor of four performance difference. It's something we still need to address. EDIT: 1 million documents / 37 seconds / 8 connections = 3378 ops/second, which is about a factor of four off from 800 ops/second that RethinkDB does on this. (This is obviously back of the napkin math.)
There's lots of work to be done, I'm very much looking forward to publishing authoritative, scientific comparison benchmarks.
I expect that if I move to binary format for inserting it might be even faster. I'm basing this since my reader is binary and that made x3-x4 improvements for my data set/computer/etc.
http://www.rethinkdb.com/docs/advanced-faq/#what-happens-whe...
Edit: Packages have been uploaded to Launchpad, waiting in queue to build. (12PM PST)
Overall, an exciting release. Going to upgrade and see what insert speed is like on my setup.
Unfortunately when migrating from 1.6.2 to 1.7.1 I lost a table (not sure how) and all secondary indexes :(
When we use a benchmark that doesn't bottleneck on latency (by adding more concurrent clients, or by using noreply) the ops throughput approaches theoretical IOPS throughput of the SSD.
Please keep in mind that this is not an authoritative comparison and it may contain mistakes. Plus as for many such systems, the aspects covered are in reality not that easy to be described in just a few words.
Platforms:
- RethinkDB: Linux, OS X
- CouchDB: where Erlang VM is supported
Data model: - both JSON
Data access:
- RethinkDB: Unified chainable dynamic query language
- CouchDB: key-value, incremental map/reduce
Javascript integration:
- RethinkDB: V8 engine; JS expressions are allowed pretty much anywhere in the RQL
- CouchDB: Spindermonkey (?); incremental map/reduce, views are JS-based
Access languages:
- RethinkDB: Protocol Buffers
- CouchDB: HTTP
Indexing:
- RethinkDB: Multiple types of indexes (primary key, compound, secondary, arbitrarily computed)
- CouchDB: incremental indexes based on view functions
Sharding:
- RethinkDB: Guided range-based sharding (supervised/guided/advised/trained)
- CouchDB: -
Replication:
- RethinkDB - sync and async replication
- CouchDB - bi-directional replication can be set between multiple CouchDB servers
Multi-datancenter:
- RethinkDB - Multiple DC support with per-datacenter replication and write acknowledgements
- CouchDB - (?)
MapReduce:
- RethinkDB: Multiple MapReduce functions executing ReQL or Javascript operations
- CouchDB: views are map/reduce but they need to be pre-defined
Consistency model:
- RethinkDB: Immediate/strong consistency with support for out of date reads
- CouchDB: http://guide.couchdb.org/draft/consistency.html
Atomicity:
- both document level
Durability:
- both durable
Storage engine:
- RethinkDB: Log-structured B-tree serialization with incremental, fully concurrent garbage compactor
- CouchDB: B-tree
Query distribution engine:
- RethinkDB: Transparent routing, distributed and parallelized
- CouchDB: none
Caching engine:
- RethinkDB: Custom per-table configurable B-tree aware caching
- CouchDB: none (?)
[1] http://rethinkdb.com/docs/pragmatic-faq/#how-do-i-take-advan...
I think that's correct.