MongoDB's bad reputation is well-deserved and based on a long history of misleading (or outright false) marketing claims coupled with poor technical properties (from unavoidable data loss to horrible compaction performance to global write locks). At this point, even if all the known issues were fixed (and they aren't) it would still take a long time for their reputation to recover. Given that some of MongoDB's direct competitors do not suffer from these problems and can largely function as drop-in replacements, recommending people use something else is really the only responsible course of action.
Which competitors are you referring to? Personally I find the power and flexibility of the MongoDB query language to be a big positive that for many cases outweighs performance and reliability issues. It seems like no other competitors are quite there.
(ToroDB developer here)
You should not base your decision of database (or anything else for that matter) on marketing copy. For something as important as your primary data store, you should at minimum read the full documentation and run some tests with dummy data to see if it will even plausibly work for your use case.
I used MongoDB successfully for years with a large data set (>1TB) and 100% production uptime for more than 3 years. I never lost data. Your claim that you will unavoidably lose data is baseless and without merit. In fact, every issue you listed has been fixed, again, counter to your claim.
Personally, these days I prefer TokuMX if I'm looking for something compatible with MongoDB, but these baseless attacks on MongoDB have to stop.
EDIT: Every time I make a post like this, I get some downvotes without responses. Please tell me why I'm wrong. If it's just that I'm abrasive... Well, you would be too if you were addressing the same thing for the Nth time.
The fact your anecdotal evidence is that you did not lose any data doesn't mean the internet is not full of people who have lost data with Mongo. I have no idea what your workload is, but my experience with data loss and uptime has not been as great as yours.
I'm not for bashing things either - I think there are cases Mongo might be appropriate, I just don't like countering claims with "it worked for me on this one data set". If it drops writes for one out of 100 people that's still a big reason to avoid it if that's a big concern for you. As for "these issues have been fixed" you're welcome to open the issue tracker - no one at Mongo claims all of these issues have been fixed (then again, PostgreSQL has open issues too) so your claim that "these issues have been fixed" is kind of odd...
And it's a little disingenuous to point at the issue tracker -- as you say, everyone has open issues. The specific things that are mentioned though have been fixed: writes are checked by default now, the global lock has been broken up into per table locks, etc. There may still be common issues that aren't being addressed, but if there are, I'm not aware of them.
Also, if I have to choose a datastore, although marketing shouldn't be important, funding is and MongoDB has had some huge funding rounds in the past. This gives me a lot more confidence in choosing it.
But I'd expect nothing less from HN.
In the face of actual data, and reproducible tests - isn't saying something like "well it didn't happen to me" dense?
The comparison might be insensitive, so excuse me for hurting your feelings, but I don't see how its invalid.
A more apt analogy then would be someone saying "My database runs with 100% uptime for 3 years, so there is no reason for me to keep backups"
We actually evaluated TokuMX extensively last year, pre pluggable storage engine. We might have pursued it if they had implemented a compatible oplog at the time, but with a migration path that consisted of dumping and importing production data -- and no way to re-elect the old primary if there were any production problems -- that made it simply a non-starter for me.
They did eventually implement a compatible oplog, which was a good product decision, but the entire TokuMX engineering team recently quit Tokutek en masse so it's still not a great option in my book. Too bad.
I would like to know how it happened. Do you have a link to an article or any other source of info?
The issue is that they even lie in their documentation[1].
Also Mongo not necessarily loses data in a catastrophic way, you might have some old or inconsistent data here and there. If you have an authoritative source of data I would compare the data in Mongo with it. Also [1] shows how you can get data corruption even with highest safety settings due to broken design.
[1] https://aphyr.com/posts/322-call-me-maybe-mongodb-stale-read...
I'm pretty sure that I could kill 0.01% of writes of any random application (which doesn't require extensive audits, like banking for example) and nobody would notice for a really long time. And if the effect was ever noticed, application code would be the first place to look at for the reason.
But there are some genuinely amazing things about MongoDB and some places where it really shines. Like flexibility -- we run over half a million different apps and therefore half a million different workloads on Mongo. Schema changes and online indexing are painless. Elections are pretty solid. And the data interface layer really can't be beat.
With the pluggable storage engine API, mongo is really growing up and becoming a real database, much like mysql did many years ago when it graduated from MyISAM to InnoDB. I'm excited about the future.
(By the way, thanks for the balanced, informative, and generally great comment!)
My first experience with it was for a social game running off of 3 mongo servers in a replica set & here at a large healthcare company we use it for several internal CRUD applications (again in 3-server replica sets) which we continue to iterate. In these use cases, Mongo's flexibility and ease of making changes trumped it's shortcomings. My understanding is that many social games still use Mongo to store player data as our studio was told to use it by a large publisher.
My take is that it's good for building prototypes and it's pretty flexible to both change & migrate data in and out & Javascript devs pick it up relatively quickly. But I also think one should be well aware of it's shortcomings and be careful not to use it where those things are of importance. More often than anything else, I recommend PostgreSQL when other people ask for a general-use db but I'd likely would pick up Mongo myself simply for speed of development.
Lastly, being part of the "MEAN stack," we can point junior devs & interns to one of the many books that cover the stack & they can learn best practices get up to speed quickly. There's an advantage to having books & a slew of SO answers to refer to in that other devs aren't pulled off of their work to teach. We literally had interns committing code on Day 1 of their jobs last summer as they came in having read up on our stack.
Regarding claims MongoDB has made, which I've only been made aware of the last two days, I don't think there's any defending that & it makes my itch to checkout rethinkdb a bit stronger than it was 2 days ago.
Knowing the limitations and behavior of Mongo can go a long way in avoiding some of the issues people have encountered.
Definitely looking forward to testing WT and RocksDB in 3.0, beyond performance improvements will drop our storage costs with compression and for a indie studio every dollar counts!
And I would disagree that Mongo is a "toy". We use it as part of our core EDW and we have petabytes of data in our data lake. No issues what so ever.
One that you don't even realize until you compare the data with canonical source.