Newsblur runs into MongoDB bug – can't replicate DB
jira.mongodb.org
jira.mongodb.org
Every last tool I have ever worked with has trade-offs. I don't have any problem with that; I've sometimes even gone as far as phrasing it as, "If you don't hate your tech stack, you aren't really using it." I can tell you all kinds of things that C#, the .NET CLR, IIS, Apache, Mercurial, Git, elasticsearch, Redis, Gunicorn, Python, Celery, and SQL Server do that make me livid, because I've used them heavily. But that's because I use them heavily. I've never had any of these tools bite me in the butt early in the process, and definitely not had any of them bite me in the butt in unexpected ways down the road. They bite me when I push them incredibly hard, right to their limits, and, due to well-understood design constraints that I'm frequently anticipating hitting ahead of time, they fall down. That's normal and fine, and handling those situations is just good software engineering. Your tools will have limits, and that stinks, but handling those limits is part of what your job entails, and you need to deal with it.
I do not use MongoDB. But here's what I see: about once a month, I come across an article where something incredibly fundamental to Mongo does not work properly. Not only does it not work properly: the way it doesn't work properly is exceedingly bad. In this case, Newsblur can't shard, which removes one of Mongo's best benefits, and the way it fails isn't to tell you early on that you will not go to space today, but rather to segfault and die after six hours of replication.
That's not predictable. That's not documented. And that's not something you can anticipate. As a developer, that concerns me, and it should concern you, too.
I understand that 10gen is an awesome, responsive company, and they have always been there to help. I don't want to malign that. When the Trello team had Mongo-related issues the other week, they were trying to help them out, too. But I genuinely do not view as paranoia my belief that the frequency and severity of stories like this mean that MongoDB is still not a good technology choice.
Given that you're (a) not using it, (b) making judgements based on a series of blog posts, (c) extrapolating bugs that only manifest under specific scenarios, (d) are clearly new to technologies like Apache, ElasticSearch, Git, Redis etc so haven't seen some equally disturbing bugs.
There are two things it's good at: 1) indexing arbitrary json and 2) searching for geo points offline.
For example, for their sharding setup they require three separate servers to host shard config data. The idea is to have high availability of that data should one of those config servers go away.
However, by design, losing a single config server can cause multiple shard masters to die, requiring a manual restart. How they die and when is random, determined by where in a block migration a server is.
How is this high availability? All three config servers must be running at all times without interruption for your cluster to be stable. Their roadmap to deal with this issue is sometime next year.
One month not too long ago, my company found 90% of the bugs listed in their bug tracker for a specific release of MongoDB - many of which would have been found with a basic unit testing suite and some minor load testing. We were effectively performing QA functions for 10gen in our production environment.
I've gone through nine releases of their PHP driver to deal with broken Data Center Awareness and none of them have worked - DCA still eludes them two years later. We have to do an OS level hack to make this work, that breaks other HA functions.
Finally, their mmap design means that memory use is extremely inefficient - on a box with 256GB of memory, with a database that is only 100GB in size, it still hits disk on db reads because they offload memory management to the OS. Any other enterprise-level DB would preload the entire dataset in memory if there's room, but not MongoDB.
It really is terrible.
Or rather: It hits the OS's disk cache. I'm not saying that this isn't a problem, but it's far from as bad you make it sound.
(http://wiki.postgresql.org/wiki/Tuning_Your_PostgreSQL_Serve... actually recommends limiting `shared_buffers` (PostgreSQL's in-memory cache) to let the OS disk cache do its magic.)a
I would love it if I could choose which server to sync from. This option used to exist, but they removed it. But that would only solve the performance penalty of replication.
To their credit, the MongoDB folks have been stellar. I had a hardware failure a year ago and the CTO ssh'ed into my machines to figure out what was wrong. This time I'm having a bit more difficulty getting the problem fixed, but it's understandable as I have 100GB of highly variable data.
Not to sound rude, but what is so difficult about managing 100GB of data? You can fit that in RAM without much work.
As for the 100GB, I have 12 task servers with 6-8 processes each reading 25-50 stories each every 4 seconds. I also have 18 app servers with 6-8 processes each reading 12 stories every second. Each story is on average 4KB. That's several MB a sec.
(Speaking as a user of Postgres and MongoDB, I'm way more worried about the latter screwing me than the former.)
This is a JIRA. It became a story once you posted it and commented on it in Hacker News.
Just waiting for the inevitable "this would never happen with PostreSQL" comments to appear.
Many of us here just don't have time to waste on technology that purports to be suitable for high-end work, but then fails us time and time again. As this incident shows, MongoDB can, and often does, end up in this position of failure.
Anyone faced with a MongoDB failure of some sort shouldn't put up with it, and should be very vocal about it. It's important to let others know the problems that can be experienced when using MongoDB.
There appears to be zero CPU or memory optimization in place. Mostly I blame the mmapped files.
True, and thanks to whoever edited the title.
Though the error in that bug was due to a regular replication event ("Fatal Assertion 16360") and this one was happening during an initial sync ("Fatal Assertion 16361"). Maybe the bug was fixed in one and not the other.