Our take on RethinkDB vs. MongoDB
rethinkdb.com
rethinkdb.com
1. MongoDB has massive storage overhead per field due to the BSON format. Even if you use single character field names, you're still looking at space wasted on null terminators. 32bit fixed length int32s also bloat your storage use. We solve this by serializing our objects as binary blobs into the DB, and only using extra fields when we need an index.
2. In Mongo, the entire DB eventually gets paged into memory and relies on the OS paging system which murders performance. For a humongous DB, not so much.
3. #1 and #2 force #3, which is sharding. MongoDB requires deploying a "config cluster" - 3 additional instances to manage sharding (annoying that the nodes themselves cannot manage this, and expensive from an ops/cost standpoint).
What I would like to know is:
1. What is the storage overhead per field of a document in RethinkDB? If it's greater than 1 byte, I'm wary.
2. Where is the .Net driver?
2. No ETA yet, but we're about to publish an updated, better document, better architected client-driver to server API spec, so we'll be seeing many more drivers soon.
http://docs.mongodb.org/manual/reference/command/collMod/#us...
10gen have been thinking about compression but nothing specific has happened yet (https://jira.mongodb.org/browse/SERVER-164). ZFS + compression is interesting, but not 'production' quality if you're using linux, and last time I tried to get MongoDB running on Solaris I gave up...
https://jira.mongodb.org/browse/SERVER-863
The issue has been open for over two and a half years, is one of the most highly voted issues, and has yet to even have reached active engineering status.
Agree with you that compression is just a workaround for the awful BSON format.
I used Mongo before and it is fine db and I don't think I would be sad to use it, however rethink really does so many things better.
Again I just started using it and things are really good, I didn't ran into any obvious limitations and annoyances.
There are several features that I really like, for example: web admin is really well done, it is easy and obvious how you create cluster, there are a lot of small things that made me jumpstart my development faster, as I can run queries in admin to try them out and I also get data back to see how things will look like.
The only thing I am somewhat missing is 'brew install rethinkdb'
+1 for sure for this one
btw, when I was installing it, I tried that even before checking on homepage :)
And this idea that joins is a requirement for a "serious" database makes absolutely no sense. Database level joins are toxic for scalability and IMHO should always be done in the application layer.
I've seen too many poor re-implementations of relational database functionality in the application layer to ever recommend it as a standard starting point. Doesn't the concept of not prematurely optimizing apply here? Solve the scalability problem when you need to. That may mean moving some join functionality into the application layer, but the solution to any given scalability problem depends on the specifics of the problem. Just throwing out database joins as a rule seems drastic.
First party MVCC is the only one that matters. It affects vital things like backups, analytical queries and transactions.
Joins are extremely useful. If a database does the sharding, it is almost always better for it to do the joins as well. Performance can be good with the right model, and Mongo is slow anyway.
As for MongoDB performance well making a blanket statement is pretty silly. On a previous project I had queries that were upwards of 40x faster in MongoDB than MySQL. Why ? Because MongoDB allows the ability to embed documents within other documents to the point where I could make a single query with zero joins to fetch 20 entities worth of data.
Every database is optimal for different use cases.
If you want a durable write; you should not disable journalling and use safe mode / getlasterror with the desired writeconcern setting - http://docs.mongodb.org/manual/reference/command/getLastErro...
Sure. Which is the default approach of almost all of the drivers.
Also safe=true only makes sure the server acknowledged your write; writeconcern allows you to wait for it to be written to the journal or more.
Also, journalling is not controllable via client drivers, only via startup flags / config options.
Not having database level joins is toxic for scalability for so many reasons.
MongoDB reminds me of MySQL : The Early Years. When every ignorant design decision and missing functionality was somehow actually a benefit. Then it gained them and most nervously smiled and moved on.
Tell that to Teradata.
However, does anyone have any practical real-world experience using it? It's not production ready (from what I gather), but has anybody actually used it for real world stuff?
For my own part, I tried it out, and got stuck trying to implement a many-to-many style join. I did some searching, and it looks like that is not really possible at this point. Not a bit deal, but it might be handy to have some example SQL-to-RethinkDB queries, just to help us newbies figure out the ropes.
There are people that told us they are experimenting with it for real apps, they've sent extensive feedback (most of these can be seen on GitHub), and some have started to build libraries for RethinkDB. As with anything young and open source, it's difficult though to tell with certainty how many projects are using a tool and what stage are they.
> it might be handy to have some example SQL-to-RethinkDB queries, just to help us newbies figure out the ropes.
Working on it already.
alex @ rethinkdb
What did you try to do with a many-to-many join? We could help you with writing the query, and could add syntax sugar to the language to make it easier if it makes sense.
I hear you, and, FWIW, I'm excited about Rethink. To rephrase my question/observation: your article clearly lays out why you think it is better than MongoDB, using some quotes from people who agree with you. However, without some real-world data, it is still an argument rooted in theory. I like theory, but I also like to take real-world data to my bosses. Do you have any stats/examples that actual compare and demonstrate the performance? (I understand that wasn't the purpose of your article, just asking as a follow up).
Regarding the many-to-many joins: I was just playing around with a contrived example: "a blog post has and belongs to many categories." Mainly I was just curious how to do it, I didn't _need_ it for anything. But, I couldn't figure out how to write it with the query language. I was using Python DSL, FWIW.
As for many-to-many joins, I'll write something up about it, thanks!!
I know they're just trying to contrast Riak and Cassandra with Couch and Mongo, and that Riak is designed to shard easily without the developer having to think about it.
That philosophy actually is "developer-oriented" in that it SEEMS like an operational savings because it was designed by developers.
Saying Riak is categorically non-operations-oriented is a bit hyperbolic, but I will be the first to acknowledge that we need even more visibility into failure-recovery / degraded mode situations. I've spoken to a few customers who have "cheat sheets" of Erlang console commands they use to debug things like handoff slowness or poor performance in general. This alone means we need to do better,
On the other hand, Riak continues to function in scenarios where other databases would be completely unavailable. I'll take immature visibility during those situations over complete unavailability any day,
I appreciate your feedback - I can assure you that this is something we're constantly working on and you'll see improvements with each release.
Finally, if you've been bitten by anything specific you'd like to see fixed, we do all our development in the open at http://github.com/basho, so github issues, pull requests, etc go right into our internal tools and workflows.
Cheers,
Andy Gross
Let's grab beers at Erlang Factory!
Chad
Also, I'm giving a talk on it in 9 days at ErlangDC...Whisper is now top 10 social networking apps and we had a number of critical Riak failures. i'll be elaborating on them, though the focus of the talk is not to bash Riak, no pun intended... just to provide our experience and how we worked around it.
Definitely looks interesting though, and I look forward to playing around with it at some point.
"Some key features like secondary indexes and live backup are still in development"
[1] https://github.com/rethinkdb/rethinkdb/issues/88
[2] https://github.com/rethinkdb/rethinkdb/tree/jd_secondary_ind...
How can I understand the performance of slow queries? Understanding query performance currently requires a pretty deep understanding of the system. For the moment, the easiest way to get an idea of why your query isn't performing well is to ask us.
Wish RethinkDB was a little further along because it seems like it might be a good fit for a new service I'm building.
We are building a tool to explain in a nice way how the query is executed, what are the bottlenecks etc. It should make it for 1.5.
You can track progress here https://github.com/rethinkdb/rethinkdb/issues/175 (it's kind of empty for now)
Currently we're using Hive and Python over streaming Hadoop. There's no significant ongoing data accumulation; we're just analyzing the data we have.
P(biased|personal) -> 1
I'm going to benchmark it when I get some time!
:-)
Mongo has been around for years, and it still has problems.
Rethinkdb is just launching a new product that essentially does the same thing as Mongo, but is maybe just a little easier to use.
I think the Yet Another Database (YAD) question still hasn't been answered by this post.
Does anyone here know any big website/service which uses RethinkDB?
«An asynchronous, event-driven architecture based on
highly optimized coroutine code scales across multiple
cores and processors, network cards, and storage systems.»
It may be a dumb question, but isn't this statement a bit contradictory? As far as I understand, event-driven design and coroutines (i.e. cooperative multitasking, lighweight threads, etc.) are the techniques usually chosen to AVOID concurrency.How does such a design imply multicore scalability? Obviously, coroutines and event loops don't prevent you from running in multiple cores. I just fail to see the correlation.
If instead we used threads + locking like traditional systems, we'd have to deal with "hot locks" that block out entire cores. Effectively we solved this problem once and for all, while systems that use threads + locks (like the linux kernel) have to continuously solve it by making sure locks are extremely granular.
Also where is the .Net driver?