Why everyone hates MongoDB
gustavoveloso.posterous.com
gustavoveloso.posterous.com
There are better (more mature, easier to maintain, faster, etc.) alternatives for both single-machine (Postgres, MySQL) and distributed (Cassandra, HBase) usage.
It's a pain to take care of, it does table scans all the time (partly because it has no schema, partly because of design issues), it has pretty poor profiling tools. It's almost impossible to use without a model/schema in the app. Migrations are a pain.
The one thing it's particularly well suited at is indexing and searching JSON-like data.
I'm surprised that MySQL Cluster Edition doesn't get more attention in the scale-out space. It is a simply brilliant product.
OMG! What a surprise!
The fact that I read the documentation and through testing discovered its a perfect fit for our application means I am not one of the 'everybody' that hates it.
I find especially during prototyping having a clear data model is great. SQLAlchemy + Postgres, then later on Alembic for migrations is so much more pleasant, predictable and performant.
My point was that it is designed for being distributed and it doesn't hide that fact. Mongo's replica sets and sharding are much less transparent and much less predictable.
When DHH created Rails, was it his job to tell you all the reasons why you should not use it over Java/.NET/PHP or, instead, to tell you all the reasons that he created it and the problems that he was trying to solve by developing apps with it?
To say that there are better alternatives implies that you understand everyone's use case. It is faulty logic. One of the reasons for the revolt that was "NoSQL" was an attempt to get people to stop shoe-horning data into MySQL ... so the concept of developers using database engines poorly isn't new. Same with MongoDB and developers that use it based on their experiences with MySQL or other relational (or non-relational) models.
I have seen people use MongoDB to great (even amazing) success and have seen others use it very poorly. In both cases, it was the developer (and not the technology) that was responsible for the way they used the technology.
Also, while MongoDB does have its faults and some sharp edges, it is far from poor. For the most part, 2.x MongoDB is solid, performs well and generally without issue.
They claim sharding works up to a high number of nodes, but even when following their recommendations, it's pretty easy to get inconsistencies. Also, many useful commands don't work (or work differently) with sharding. This is the sort of thing Cassandra and HBase get right.
They used to claim you can have control over whether reads/writes happen on the master or slaves, but not all commands supported this differentiation. There's was safe and masterOnly, but they don't always apply (and amazingly enough, don't always work). This is the sort of thing Cassandra gets right, and at least newer Mongo (and clients) support more granularity, and more consistently.
They claim certain benchmark numbers, but neglect to mention they're done without durability or any router/replica/shard setup.
It's nothing overly enthusiastic creators of a new database wouldn't be expected to do, but still annoying.
The same thing happened with MySQL and now it looks like it is happening with MongoDB as well. Mongo has so many corner cases and although I have never scaled it up in production the amount of moving pieces and the clear lack of attention paid to distributed programming literature in the design doesn't inspire a lot of confidence.
Technical problems with MongoDB are being explained away as features in the same way MySQL did. Why does Mongo MMAP its datafiles?? So the cache stays warm!!! Why doesn't MySQL have transactions? No one needs transactions, just lock the table !!
If you didn't consider these cases you might think that Mongo sounded pretty good. I don't fault 10gen for pushing this hard and taking advantage of the perception that NoSQL is a more modern technology. I just can't believe there isn't something better that does what Mongo does.
"'And in conclusion, we have found MySQL to be an excellent database for our website. Any questions?'
Yes, I have a question. Why didn't you use MongoDB? MongoDB is a web scale database, and doesn't use SQL or JOINs, so it's high-performance. ...."
Read the entire conversation at http://www.mongodb-is-web-scale.com/
And if you just chose it willy-nilly because it's a hot tech, well.. You made a mistake.
PS - linking to a couple paragraph blog with come conjecture and a link to a google search? That deserves a bitch slap. Try harder.
However, we can see that it has again sparked the age old flame war. It reminds me of the online flamewars that start once someone voices their opinion on some hot topic like mac vs pc, ios vs android, etc.
Nowadays, I have begun to defer my database choice until I absolutely have to put an app into production if possible. That way I can write an app and have it working without worrying too much about queries, caching, and so on until I need to actually put it out on a real server.
MySQL, Mongo, HBase, Cassandra, Postgres, Riak, Redis, Memcache, Couchbase, Filesystem? It doesn't matter.
They're all going to break at different times for different reasons depending on how you use them and what scale you are at. Everything breaks at scale eventually.
Build something awesome enough that it breaks under the user load. That's a good problem to have.
When a product is pushed as essentially a silver-bullet -- often by people who have never and will never actually build anything with it -- it is naturally going to get a lot of backlash by people who see through the problems, and moreso by those who buy the hype and then smash into the reality halfway through their projects.
I speak from personal experience on the first point, while the second has come from countless real-world case studies. A couple of years back I had several widely distributed pieces declared "trolls" because I essentially called for a sobering reset on the NoSQL hype train. Every single thing I said has been well proven out since, so there's that.
My side projects are where I have the most freedom and I just recently built and open sourced Obvious - http://obvious.retromocha.com. It makes it easy to build apps with good structure, good testing, and with pluggable front and back ends. So, if I want to switch out Mongo for MySQL or Postgres, it's basically painless.
If I had a bunch of keys and values, I'd put it in memcache. If I had relational data, I'd put it in postgres. I don't see where mongo fits into play. Every time I've seen it used in production, it always seems like its being used as an improperly implemented RDBMS or an improperly implemented key+value store.
The setup alone is just bizarre, for each replica set you have to have 3 (and only 3!) config server instances associated with your set, along with possibly an "arbiter" instance depending on what circumstances you're operating under. I could do a whole rant (and have, in the past, to anyone who would listen) on the arbiter/voting system mongo uses. We've been left in a situation in the past where one of our secondaries randomly died (more on these lovely "random occurences" later), leaving the cluster with a master and an arbiter left over. Mongo decided that since there was an even number of votes (and why is my production-critical database voting on things?!) it couldn't promote a master (even though the master never actually died), dropped the master down to being read-only, leaving us completely boned. They have since fixed this issue (I think), but it definitely garnered a lot of ill-will from me, and should never EVER have even been a problem.
"Random occurrences" have always plagued us with mongo. Whether it's been random segfaults (less common with more recent versions, but oh dear god were they frequent pre-2.0), secondaries being promoted to master with no clear reason as to what sparked this (which has left us flailing twice now, when a master decided to step down while one secondary was in RECOVERING mode and the other was AWOL), config servers getting out of sync (one time one of the config server sets decided it was going to start hosting the config for an entirely different replica set, luckily no production mongos instances had to reload the config before we noticed), mongos instances getting out of sync/crashing (less common now than before, thankfully), and I'm sure much more that I'm not thinking of now.
To make all of the above about 10x worse the mongo logs are terrible. There is no distinction between INFO messages and ERROR messages in the logs, so everything has to be treated as an error of some kind. And there are many "errors" that we should "just ignore". This makes debugging pretty much impossible. Another nice feature is when you restart a mongos instance (and possibly a mongod instance too, although I could be wrong on that) all the logs from the old process get clobbered by the new one (so backup those logs folks!). It's extremely difficult to track down why slow queries are happening unless you can catch them in the act, and even then it's not trivial.
Mongo has many great qualities. It's one of the few (if the only, that I know of) that can do many of the things that it does, and for that it's very useful. But for it to have been marketed as a production-ready database was a bit disingenuous I think. It scales OK. Not well, just OK. You can make it scale if you try real hard and tiptoe around it so you don't wake it up and make it cry. I think someday in the future all of these issues will be fixed, but at the moment I don't recommend anyone use mongo for anything they plan on a significant number (200k+ users at any given moment) of users being dependent on.
On the logging, not sure what you mean by the old logs get clobbered. Is this, perhaps, some house-cleaning job that you have on your server? When we stop/start processes, the logs are appended and keep going.
On the issues with primary/secondary and state changes, do you host your databases on Amazon? We have often found that when we have this issue, it was pinpointed back to some temporary intra-zone networking glitch with AWS ... either with their name resolution or a blip in one zone being unable to see another zone in the window of time that MongoDB has set for the response.
These often (if not all the time) go unreported by AWS until you get them to dig a bit and report back. We effectively utilize priority to keep the member we want as primary and if a state change does occur, it moves back once the full set is healthy again.
For us the logs are not opened in append-mode, just write. We don't do any house cleaning except for logrotate, but that only runs nightly and isn't the cause of what we're seeing. Like I said, it may not be for the normal mongod instance, but it's definitely true for mongos. I'll double-check on the version number of what we're running
> do you host your databases on Amazon?
Nope, all owned servers, over a local network on a switch we own as well. The switching happens on some clusters more then others, so it's possibly a usage related thing.
On the point of slow queries, nailing a slow query is a bit of a mixed bag at the moment. I've found that Dex (http://blog.mongolab.com/2012/06/introducing-dex-the-index-b...) and NewRelic help diagnose trouble queries pretty well.
http://docs.mongodb.org/manual/tutorial/manage-the-database-...
Re: replica sets. You don't need config servers to run replica sets. You need config servers to run a sharded cluster. Your understanding of when to run an arbiter is also flawed. I think you might find the documentation enlightening: http://docs.mongodb.org/manual/. The unfortunate situation you describe where the master was left in a read-only state is a flawed setup -- not an error in mongo.
In your second paragraph referring to "random occurrences" you're again conflating sharding and replica sets... I'm not sure what exactly you're describing, sorry :(!
IN general you should not ignore errors in logs... the mongo processes all have log rotation via signals or server commands.
I think a lot of what you're describing is hard to take at face value because of a lack of a common vocabulary -- or some misunderstanding. In my experience mongo has worked fairly well in some large scale systems. It's not the solution to everything -- I'll agree with that... but if you read the documentation it does what you would expect.
MongoDB stills young and received a ton of attention for its promise. And is really unfair trying to compare it with 20+ years technologies.
MongoDB users stills pay the price to be an early adopter and I'm confident we are doing a great deal and learn a lot.
Pretty much any database is going to leak abstractions at some point, there's absolutely no way of getting around understanding how your DB works. I think the problem with Mongo is that it's so developer friendly initially, that it makes it easy to avoid a lot of hard decisions that will eventually have to be made, and when those decisions do come to the forefront, it's often at a much less convenient juncture in the application's lifecycle to change things.
1. A capped collection of responses from a 3rd party endpoint (XML) that we often want to go back up to 2 weeks and have a look at.
2. Data that is all reads after initial write: Historical data, lots of it (100-150M rows and growing fast), that we want to take advantage of the scaleout nature of MongoDB to handle.
3. "Built out data": Data that is the summarized output of the ground truth in our MySQL DB that is expensive to compute, has a varying JSON schema, and we want have ready at our fingertips for the frontend.
It seems like MongoDB is well suited for these use cases, leaving our transactional and user data in MySQL.
Am I crazy? Should I be considering one of the alternatives instead?
And I really didn't understand other people complaining about 10gen marketing capabilities ;)
Besides antirez being a friend, I think something like redis more credible due to the fact that it seems fairly clear about what it is, and what it is not.