Does everyone hate MongoDB?
blog.serverdensity.com
blog.serverdensity.com
I am indifferent to MongoDB but I do caution people that the internals and implementation are quite primitive for a database engine. It is the sort of database engine that a programmer that knows little about database engines would build. Over the years that has repeatedly manifested as problems no one should ever expect a database engine to have if properly designed. There is a reasonable point where you should not have to read the documentation to see if your database has a basic design flaw; there is an assumption of competent architecture for its nominal use cases.
MongoDB has improved in this regard over time and many other popular NoSQL (and SQL) databases have similar problems. Users are being asked to cut the database some slack for less than competent engineering choices but users just want a database to work. Being "simple" isn't enough.
Agreed, though I think its also coupled with just a lack of understanding of what MongoDB does beyond just "its NoSQL!". So many people are used to working with RDBMS. They just seem to make incorrect assumptions about MongoDB based on what they thought were universal rules about databases.
Mongo's a tool. Its not right for all situations and it definitely has some maturing to go still, but its good to use in certain situations, provided you're properly informed about it.
I would only add a lack of thinking on the part of developers who jump on the NoSQL bandwagon for the buzzword without understanding the implications of not having a relational database.
I'm not a RoR fan but the ActiveRecord model is probably enough for 75% of people to get their projects started without having to dig too much.
- return from a write call silently, when the data wasn't written and will not be.
If you are going to break conventions on a multi-decade tradition, there should be warnings everywhere. Not just the Downloads page, but during installation (which is harder to ignore, as the Download page is not seen by people like me who use package managers).
The 32-bit issue is worse. There's no excuse for not warning people that the server they just started is a crippled data store limited to 2GB.
Again, I like MongoDB. Given my experience with software development, those issues are telltale signs of a product in its infancy. Hence the "10 years" title of my post, which of course is hyperbole. But it will take time for Mongo to be user-friendly enough.
By the way, those who say that a piece of software has no need to be as user-friendly as possible, and you should thoroughly read the manual before even experimenting with it simply live in a different world. The more of a head start the software can give you, the better. The fewer the gotchas, the better. That was the philosophy of IndexTank, and it worked.
Can you confirm this (because I don't have the 32-bit version installed)? It still stinks, especially because it "silently failed after hitting the threshold", but I personally would feel better about 10gen if this little story is true (i.e. they warn you about it not only in downloads page, but every time you run mongod).
Just as a data point, I installed MongoDB on 32-bit Ubuntu 12.04 from the Ubuntu repo. At no point during install or service startup does it say anything.
Installing from source may be different.
mongodb start/running, process 13660The 2GB 32-bit limit of MongoDB seems like a complete non-issue to me.
Wed Sep 19 17:29:21 [initandlisten] MongoDB starting : pid=3765 port=27017 dbpath=/var/lib/mongodb 32-bit host=deepthought
Wed Sep 19 17:29:21 [initandlisten]
Wed Sep 19 17:29:21 [initandlisten] ** NOTE: when using MongoDB 32 bit, you are limited to about 2 gigabytes of data
Wed Sep 19 17:29:21 [initandlisten] ** see http://blog.mongodb.org/post/137788967/32-bit-limitations
Wed Sep 19 17:29:21 [initandlisten] ** with --journal, the limit is lower
Wed Sep 19 17:29:21 [initandlisten]
Wed Sep 19 17:29:21 [initandlisten] db version v2.2.0, pdfile version 4.5
Wed Sep 19 17:29:21 [initandlisten] git version: f5e83eae9cfbec7fb7a071321928f00d1b0c5207
Wed Sep 19 17:29:21 [initandlisten] build info: Linux domU-12-31-39-01-70-B4 2.6.21.7-2.fc8xen #1 SMP Fri Feb 15 12:39:36 EST 2008 i686 BOOST_LIB_VERSION=1_49
Wed Sep 19 17:29:21 [initandlisten] options: { config: "/etc/ mongodb.conf", dbpath: "/var/lib/mongodb", journal: "true", logappend: "true", logpath: "/var/log/mongodb/mongodb.log" }
Wed Sep 19 17:29:21 [initandlisten] journal dir=/var/lib/mongodb/journal
Wed Sep 19 17:29:21 [initandlisten] recover : no journal files present, no recovery needed
Wed Sep 19 17:29:21 [initandlisten] waiting for connections on port 27017
Wed Sep 19 17:29:21 [websvr] admin web console waiting for connections on port 28017A shitty default, no doubt. But it can be changed easily. And "most" drivers offer that. And you usually connect to MongoDB using a driver.
Defaults usually don't get changed, convention over configuration, etc.
So it does happen.
Wait they are promoting this as a database. It is right there in black on yellow "database". I don't know about you but I suspect most people expect databases to try their hardest to protect data. That means also having safe default. As in when I install it and put data in it, by default it should try hardest to make sure that data doesn't get corrupted. If it doesn't, it doesn't deserve to call itself a database.
> A shitty default, no doubt.
This is not a "oopsie" this is a deliberate lie and misinformation in order to produce fast benchmarks.
This is a message that MUST be displayed on the console when you install the server for the first time. It's too important.
Also, you learn a tool before going into production. I never went into "production" with mongodb. All I did was experiment with a toy project. I never needed to look at the log.
When I first install something new that I'm going to rely on - yes.
I install something new. I kick the tires. It takes a long time until I decide I'm going to rely on it. I'll look at the log at some point, but not necessarily the first time I install something. There are better things to do at that point.
There are some really disrespectful people in this community.
It has async writes...this is pretty well documented by 10gen and is also something noted by a lot of tutorials, blog articles etc. You should have known something this basic about a database so important to your business.
None of that would bother me in the slightest if you were not still here defending such basic mistakes and blaming them on 10gen.
2GB limit is clearly mentioned in the logs, that's fine, but anyone that sees this would expect that DB would start "screaming" loudly wherever it can (logs, response to the user on EVERY communication with the server, during any select,insert and others) that it reached this limit.
Your article begun and ended with sarcastic remarks about the product. Realistically, what kind of response did you expect? The issues described in your article are very real, and very worthy of repeated discussion, but the article itself eschews discussion in favor of pontification, sarcasm and flamebait.
The fact that a 32-bit memory-mapped file is limited to 2GB is basic comp-sci. It is also reported by MongoDB everytime it starts. Furthermore it is noted in the 10gen docs in several places.
Async writes...also extremely well documented.
So the failure here is two things: failure to properly research and understand a technology critical to their business. And then failure to take personal responsibility for the first failure, and instead to blame the vendor for not building a tradition DBMS despite the documentation about the stark differences.
I use MongoDB in a side-project if that is important to understand.
All this time I've been talking about a toy app that I wrote to kick MongoDB's tires. If you don't understand that, there is no point in having a conversation.
I never had or have any plans to use MongoDB for any business.
~$ sudo service mongodb start
mongodb start/running, process 7680
Look guys, MongoDB uses a memory-mapped file. How can it be larger than 2GB on a 32-bit system? There should not need to be "warnings everywhere".
When comparing e.g. 32bit firefox and 64bit one, the first runs much smoother on my system. So if I didn't have 8GB I would install 32bit system. (not to mention that all binaries take less space)
> - return from a write call silently, when the data wasn't written and will not be
There isn't some defined set of rules for how a database should operate. This attitude implies that an asynchronous database should never ever exist. If that's the case, how could I ever use a database for HTTP logging? I can't have every single HTTP request block on a database write, that's absurd. HTTP logging is impossible with MySQL or PostgreSQL for exactly this reason.
Is a process that send data from socket > /dev/null a database as well then. Why not call that a database too?
> If that's the case, how could I ever use a database for HTTP logging?
In a normal database, you would possibly switch 'durable writes' and possibly 'time expiration' feature in the configuration file from the default OFF to ON.
The reason for the deprecation seems to be the many gotchas with it combined with the doubtful performance gains. It seems to have been an ugly hack.
Still they manage to get those databases up an running without any data loss with just some generic knowledge about OS:es and hardware and reading choice parts of the documentation. Strange, huh?
I am not saying those are bad, but with more established DBs like MySQL and PostgreSQL, you never really saw the same kind of marketing efforts towards developers, startups, etc. It is kind of a newer concept.
That's funny because MySQL used to do exactly the same kind of marketing fifteen years ago: comparing itself to Oracle, "it's a thousand times faster" when it obviously didn't do one thousandth of what a RDBMS does.
History repeats... and has a strange humour sense.
Actually I think that MySQL is essentially the fore-runner of NoSQL. People who have cut their teeth using MySQL as a single-app persistence store or were trying to code to lowest common denominator in order to be portable now see benefits in ditching the idea of a shared database altogether because, well, they never used a shared database.
A large part of that is that 10gen are much better at marketing than they are at tech. For example, they like to pitch against the Oracle database, and for certain use cases, MongoDB is better than Oracle RDBMS. But if you were to take those cases to Oracle, they would say well don't use the database for that, use our other product Coherence (which was Tangosol before they bought them). And Coherence spanks MongoDB in every possible way. You could write a bestselling novel about it, 50 Shards Of Grey.
Another example here http://news.ycombinator.com/item?id=4533760
And if you hate Oracle, there are a bunch of free things that do what MongoDB does a hell of a lot better. Why waste time with its cheesy MapReduce when you could have ICE http://www.zeroc.com/overview.html for example?
MongoDB is flexible and fast, but practically speaking it's a hybrid product that defies simple categorization. That means it must be well understood to leverage it. Ultimately that may also be it's Achilles heel, but in the short term, it's extremely frustrating that 10gen can't embrace this fact. Instead they perpetuate marketing that causes the product to be perceived as flawed by their target audience.
Once you've discovered this, you can use it appropriately (either by overriding the default silent failure to use it as for durable persistence, or only using it for caching or as an eventually consistent store) to reap it's flexibility and speed.
(Largely quoted from my comment yesterday, https://news.ycombinator.com/item?id=4566518.)
I have since learned that I could have changed some configuration options to make MongoDB less likely to corrupt and/or lose data. I'm still wary though - the fact that the default configuration was prone to unrecoverable data loss suggests that any time I use MongoDB, I must carefully research the feature I'm using to make sure I don't do it in a way that causes data loss.
I believe dangerous configurations should never be the default, and dangerous features should be clearly labeled. The default method for writing data should not fail silently, for example. The fast, fire-and-forget write should be called something like "unchecked_write".
If they were they would go with something like MemSQL, Redis etc.
Yes do. Everytime I tell people how Mongo can lose their data and silently corrupt it. They say "ha, but you never experienced it, how do you know". Well, apart from explaining the way write work in theory I have the "There is this one guy on the internet,... let me find his blog". But I suspect there is more than one guy.
All I can do is laugh and keep on using Postgres...
My approaches to db interfaces have been greatly inspired by REST and SOAP. The thing is giving up on the RDBMS is usually the wrong choice. If using NoSQL, usually it is best as an adjunct to the traditional RDBMS (particularly for pre- and post- processing).
We have definitely had our moments where we've screamed that we hate Mongo and are going to rip it out. That's normally where we've overlooked a detail of how it works... and we've had this same experience with every database technology we use - including MySQL and Redis.
The day after, when we've cleared our heads... we're happy again. The same as the other technologies.
I think that with all technologies you're going to get bitten by some detail you didn't know about or had forgotten. The trick is to mitigate these disasters by thinking about your failure cases.
Anyways, I'm sure a lot of us would appreciate it if you wrote about your experiences, regardless of how mundane they are.
This is such a strange thing for me to observe, but I think that it has to do with the fact that MongoDB works so smoothly initially with a default install that it hooks you in, and then later when you have a large dataset and it stops working well, it's hard to understand how what was such an amazing technology can now fail so badly.
It's also a huge problem that once you run MongoDB at scale, you desperately need experts to fix things, but it's so hard to find any so-called experts who can help. 10gen did a great job of marketing to developers, but unfortunately they seemed to spit in the face of DBAs and Ops people a long time ago by proclaiming them to be unnecessary and archaic. Running any significant MongoDB database in production requires as much expertise as someone running a big MySQL instance, but there isn't a community of database lovers around MongoDB who you can hire.
The DBAs and operations people I know dislike or even hate MongoDB, often for valid reasons that 10gen should address such as the as-yet-unfixed 'write lock' issue, code instability and inaccurate/misleading documentation. 10gen has done a great job at evangelizing to developers and making features that developers love. Now that 10gen has so much money, I hope that they can now afford to start making MongoDB a database that Ops people and DBAs can love.
As a final note, this ServerDensity guy is clearly looking at the world through mongo-colored glasses. It makes sense I suppose since he is all-in with MongoDB for his company. But we were using his service up until three months ago, and had a lot of problems with intermittent performance issues that seemed to be database-related since the site was still working and only certain pages would take 30 seconds or so to load. It's possible that the problems were temporary and no longer exist, but it brings home my point. If MongoDB experts can't make their own services 100% reliable, what hope does a regular startup have of getting MongoDB to work well at scale.
A developer may feel comfortable making the decision to go with MongoDB, but if they are wrong it won't cost them their job and they won't need to pull their hair out dealing with Ops issues all the time. If a DBA or Ops guy is being hired to manage a company's datastores, I don't see MongoDB (even 2.2) being a contender. At this point there are simply other DBs available that can perform the same or nearly the same without all the fussiness. Developers may be unhappy since nothing yet is as easy to develop on, but they'll be happier in the end when stuff 'just works'.
Developers often have confirmation bias for their choices—it's a common human fault. We do it for purchases, life decisions, and software decisions alike. But it's important to watch out for it and be aware when you might be in the thick of your own bias.
A friend of mine (let's call him "me") used Adobe Flex for a relatively large computationally-expensive project, and advocated for it because of many small details and relatively quick learning curve, and ease of UI prototyping. After we got entrenched, we started running into many shortcomings which I (er, my friend) should have realized early on were due to the nature of the platform itself, but I continued to defend Flex because it was my decision and I had put so much time into making it work.
In the end, we realized it wasn't the right fit for this specific job and moved to a more capable platform with fewer issues (aka, "any other platform than Flex").
Before making big decisions, I try to remember that one mistake, and get out of my own way first to look at it from an outside perspective. I try to differentiate when I'm making decisions based on fact, or if I'm just trying to justify patching holes in the titanic. This "looking at the world through [insert technology]-colored glasses" idea is spot on, and everyone—not just those using or debating MongoDB—should be aware of.
Read. The. Documentation.
Every single issue mentioned about MongoDB is in the documentation and pretty fundamental to how it works.
If you do the five minute quickstart guide, and read the feature list, and throw it into the mix, well...
No, it's about obfuscation vs clarity, and to an extent, about how good the overall design really is. Look at it as a measure of quality. If I have to read every word of the documentation to find the tiny part on page 18 where it tells about that one flag I need to start it up with in order to ensure that one feature works the way I want; versus a sensible default and clear documentation and even clear design where perhaps that flag isn't even required (think automated memory management that Just Works, versus a dozen command-line or config-file switches about memory buffer size and such).
When you run into issues, it's not a useful excuse to say "Well, you should have read the documentation, it was all there." That's like a shady credit card contract where the rate goes up 30% if you miss a payment. It's easy, even if you read it, to say "well I won't ever miss a payment," and even easier to miss that clause completely. Can you blame people for not reading the fine print? Sure, it was their responsibility. Will people still do it? Of course. Will people who read the contract fully still run into issues if they make one mistake? Probably.
And the point: is it a crappy contract? Yes. It has a crappy feature to begin with, and it's made worse by its obfuscation. This is almost unarguable: a credit card with a lower rate penalty is better. A piece of software with good defaults and sensible design is better.
The requirement for asinine and lengthy documentation isn't just a big warning sign that you should read it—it's a sign of poor design, or at least the word everyone's been using to describe Mongo: immature.
Good design includes the whole experience of using the software, and takes into account good integrated systems and human interaction. Bad design requires you to read the documentation extremely carefully. These are not hard and fast rules, but they're certainly warning signs. Really clear and obvious warning signs. And the overall point is that the well-documented issues that come up with Mongo are not just stupid people who don't read documentation—they are statistically relevant pieces of evidence pointing to some poor design decisions.
You should really think of this whenever you see a trend. You have a choice. You can blame individuals for "not reading the documentation"—or you can look at the systematic trend and statistically evaluate the problem. The former allows you to quickly dismiss issues on an emotional basis and make yourself feel better, while doing absolutely nothing to solve the issue. The latter lets you collect useful information and make real changes that have a real impact on the issue at hand. Your call.
WTF ? I have never heard 10gen say anything remotely like this. And it surely hasn't come across in their marketing. I mean seriously. Which developer thinks that in a production environment they are going to be the ones supporting the database ? Nobody.
> If MongoDB experts can't make their own services 100% reliable, what hope does a regular startup have of getting MongoDB to work well at scale.
The fact that you base your impression of MongoDB based on ZERO evidence just conjecture that the slowness of their site is database related says a lot about you too. There are many other reasons it could equally be: app server, network etc.
> If a DBA or Ops guy is being hired to manage a company's datastores, I don't see MongoDB (even 2.2) being a contender.
Have you even worked at an enterprise company before ? DBA/Ops aren't the ones in control. If the development team wants MongoDB installed and have business justification it gets installed.
> Developers may be unhappy since nothing yet is as easy to develop on, but they'll be happier in the end when stuff 'just works'.
But it does 'just work' that's the whole point. Developers aren't stupid and MongoDB is not the only database around. You just need to (god forbid) understand how the thing works.
>>WTF ? I have never heard 10gen say anything remotely like this.
From the "Ease of use" section of their Philosophy page http://www.mongodb.org/display/DOCS/Philosophy,
This means that MongoDB works right out of the box, and you can dive right into developing your application, instead of spending a lot of time fine-tuning obscure database configurations.
There goes the indirect jab to the RDBMS DBAs wasting their time.Do you need a DBA to get MySQL running ? No. Oracle ? No. SQL Server ? No. That's all it means. Normal people understand that there is a difference between getting something running and deploying it into production.
Those other things are trivial to fix (e.g. add more app servers) so I too think that it's safe to assume it's probably the data store's fault.
> But it does 'just work' that's the whole point.
Did you read the article?
Adding more app servers increases the number of requests you can handle. It doesn't make slow apps or latency faster. Both of those are possible causes for why certain requests may be slower. It's not necessarily a database problem.
> Did you read the article?
Yes. I read the documentation before installing MongoDB so I haven't had any problems (so far).
Headline Problem: Too many people are ranting about MongoDB.
Mistake: Too few data points. Overly defensive of MongoDB.
Comments: You should probably learn more about Riak. The "From MongoDB to Riak" is light on details but isn't really hand-waving. The author is saying that having a masterless database results in less fretting. http://wiki.basho.com/Riak-Compared-to-MongoDB.html
Also, why two rage faces? Can't you get your point across without resorting to silly memes?
(in reference to I’ll Give MongoDB Another Try. In Ten Years, http://diegobasch.com/ill-give-mongodb-another-try-in-ten-ye...)
I read the referred article and the HN discussion...but wasn't quite able to tell how much of the debate was focused on the reasonability of this default rather than "developers should read the docs, or accept catastrophic failure" attitude (which is not inherently wrong).
What immediately comes to mind is the Rails ActiveRecord update_attributes vulnerability...the default was to allow the updating of all specified attributes with the assumption that no competent developer would trust unsanitized input from the browser. After a good Samaritan performed a spectacular hack on Github, the Rails team immediately changed the default.
Is that the situation here with the 32-bit silent fail default? That it is a sensible default, but could be changed if it's shown that competent devs will nonetheless screw it up?
It's kind of sad that we'd need an example to show that it really could happen in every single specific case. It should be common knowledge by now that competent devs screw things up all the time.
For example, I'm sure the MongoDB devs are extremely competent. But all the same, having a database management system default to letting writes fail silently is a pretty spectacular screw-up. I can't really blame other competent devs for taking it for granted that a DBMS wouldn't do something like that.
"Mongo sucks f*cking balls" is hating it. "Mongo silently does not store data in some situations" is pointing out a flaw.
It is easy to dismiss valid criticism as hate. It is also unfortunately true that some topics degenerate to flamewars on HN. Why Mongo is such catalyst is a phenomenon in itself. One that I find fascinating.
I am a little bothered when I see a working set described as though it was a property of the user or workload.
A large part of database research has gone into allowing the user or the system to make the working set smaller. An obvious example is an index, which makes the working set smaller if you don't mind a random I/O or two per lookup (of course, that only works for indexable queries).
There are also many operations in a database which try to work within a limited amount of memory, and therefore must have a small working set regardless of the data size. Sort and HashJoin are two examples. HashJoin doesn't tell you what the working set of your data is, you tell it and it works as efficiently as it can in that amount of memory.
And you can design your data layout to have a smaller working set (again, so long as you allow a few disk accesses outside the working set). Normalization and vertical partitioning (i.e. splitting a table up into several tables with fewer columns each) can help here.
So, the "working set" isn't some passive constant that can't be managed.
I was not criticizing mongo per se, I was criticizing this post and other misguided statements about the working set as it relates to database software.
MongoDB "just works" out of the box for testing and development. It's practically the perfect database for any testing and development work. And their marketing documentation tells you that it'll just work in production and will scale really well. That makes a lot of people try MongoDB, and they're mostly very happy with it.
...the problems arise later. Because while MongoDB lives up to the hype for testing and development, it doesn't completely live up to the hype for production uses. It doesn't always scale that well. It isn't always as durable as you might hope. Performance isn't always very good. For some use cases, the default settings aren't appropriate; for other use cases the database design itself isn't appropriate. Backups and replication aren't quite as easy as you'd hope.
In short, the problem comes because MongoDB is as good or better than, say, Postgres or Riak during development in basically every way, but it's only better than Postgres or Riak in a few specialized ways during production.
Example: Riak basically won't work at all in a multi-tenant setup. Two competing services will provide you with a free development MongoDB instance on Heroku, but nothing of the kind exists for Riak; you basically need to roll your own cluster. That makes MongoDB really easy for someone thinking "hey, I wonder if Mongo would work for this proof of concept I'm working on" (hint: it will). Of course, MongoDB replication and scaling is honestly mediocre, while Riak's is basically magic. This makes Riak a better choice for someone who needs to scale a cluster of database servers - but you only find out that sort of thing first-hand during production. At which point, you make an angry blog post that hits Hacker News. And while plenty of people never run into MongoDBs weak areas, they never post anything that hits the Hacker News frontpage. :)
TL;DR: A lot of people hate MongoDB because it's amazingly easy to get started with, but harder later on, and they then feel betrayed. Other databases have smoother learning curves.
Basically, they had chosen MongoDB due to the 'cool/new' factor, and it came back to bite them on the ass.
If you can't be assured of structure of data input, you can't reliably and gracefully transform that data on output. That's a pretty fundamental tradeoff between SQL and NoSQL (and historically between PostgreSQL and MySQL too).
It seems to me that there are three camps of people when it comes to dbs - (1) The "I hate anything new" camp, (2) The "I hate anything old" camp, and finally (3) The "I will pick the best tool suited to my needs"
Those from camp (1) are the ones who hate anything that IS NOT relational or sql. The mongodb bashers would fall in this camp.
Those from camp (2) are those who hate that IS relational or sql. The "web-scale" people would flal in this camp.
Those from camp (3) are the silent majority who do their independent research, pick the right type of DB for their job, and live life happily.
For my startup (Semantics3), we deal primarily with JSON strings (which have no fixed structure). We use MongoDB because a really good use case for it is to just store JSON documents with a unique id. We only run simple queries on it - basically existential (does an id exist) or count.
We are aware that it's query performance is not as fast as that of a relational db. So we then index that data to ElasticSearch and run our advanced queries through that. We are really happy with our current system and it has been working great.
If we had some sort of "relational" data with a fixed structure I sure has hell would have picked MySQL or PostgreSQL and used Sphinx on it for indexing.
Peace.
Those that see through the lies are probably laughing at them, those that already bought the product and found out it is useless and doesn't live up to touted capabilities probably hate them.
Now don't get me wrong. In a practical way (disregarding ethical consideration) they took a risk. They quickly gained a lot of customers because their benchmarks looked very good.
So in a certain way people probably already forgot about the competitors at that time, because those competitors never made it so far. Now MongoDB is here years later and stories of hidden data corruption (and this is one of worst thing that could happen to a database) have started to emerge. Yes, by this time Mongo has the mind-share but now it is time to also pay the price for their design choice. They are probably still ahead.
If they're using a single-thread JS engine, can't they just multiprocess? No shared state, after all...
I'm glad that there are proponents of MongoDB like ServerDensity, it's nice to hear the other side of things from people who understand the issues.
Out of all the issues mentioned, I'm particularly interested in a fix for uncompressed field names.
To be honest, I'd rather have on-the-fly compression of data fields. At the moment I need to do this on my end, but it would be nice to be able to mark fields as compressible. Ah well.
http://blog.serverdensity.com/removing-memcached-because-its...
NoSQL, in many ways with MongoDB leading the charge, came on to the scene very quickly with a lot of support on HN. When action outpaces education, then it's natural to expect education to play catch-up for a while, which is what's happening now.
Some of it is education by 10gen, some by users who made mistakes (the leading cause of education), and some by traditional database people who want to highlight the lessons learned that might have been missed by 10gen or mongo users.
Generalizing all of this as "hate" is not productive.
Sure it is. People download MongoDB, run some simple workload against it, and see how fast it is. It's not an official benchmark, but it's still a strong part of the marketing message.
So, if those people don't understand the significance of sync/async, they may be getting a false impression of the speed.
The presentation of mitigation strategies for both sources of issues was actually constructive for potential and current users of Mongo. As with any tool, Mongo is useful in some scenarios but not without its drawbacks. It's nice to have a point-by-point assessment of pain points.
Falling into the potential user category, I've opted not to use MongoDB for my current project because the particular benefits don't align nicely with my problem, but it was useful to see this and know a bit of what to expect with Mongo, good and bad.
I don't hate MongoDB but I think it is overhyped. The fact that 10gen raised 42m is crazy. I would be surprised if those investors got any return greater than the money they invested.
With anything there are gotchas so you have to live and learn with em.
When I was researching Mongodb before going into it I saw forums talking about the 2gb limit on 32 bit systems; I would say the author of that article didn't see the posting - which could either be a case of not due diligence on the ins and out of the system or lack of diligence on the reference material he had originally been using to build his system.
I have an event logger, that logs "interesting" MySQL server events to a tab separated file and then out to a central mongodb database. So I can query several servers results at once.
Yes with 2.2, write locks have improved drastically. However, I'm still waiting until you can get to the point of not having to lock down whole collections (the near equivalent to a table - with caveats) as opposed to a document within the collection (the closest thing that you can call a record in db terms - with tons of caveats).
This being said, Mongodb does have really really awesome documentation and it is the right way to go for many applications and developers.
However, I am now very aware that I have an implicit schema in my code that will be harder to understand in 6 months than a database schema.
I traded fantastic flexibility in initial development against easily maintained longevity.
I suspect I will move to something more relational but have also taken on the need to really understand the trade-offs I'm making. I've started reading Seven Databases in Seven Weeks to fill in some of my gaps.
Great quote, btw, and spot on.
I especially like MongoDB for data analytics: really handy for storing data and then doing experimental data mining. Putting a read replica on each server that does analytics/data mining works especially well (very good performance when MongoDB's indices fit in memory and the read mongo is on localhost to programs doing analytics).
For a regular web-app with less exotic needs, my store of choice is PostgreSQL these days.
And next time I have to handle large scale data, I will probably use Riak.
Thank you for going back through mongo articles to iterate the key failures to understanding points.
site:jira.mongodb.org/browse/ planned but not scheduled
However, if you just want to write a web app that doesn't lose data it's a really crappy idea and should probably use something creating for solving that problem, like a compiler and web framework.
Honest question. Do you actually leverage the schemaless nature of MongoDB? Not 'worrying' about managing schemas is one thing, but actually leveraging the fact that your objects can have different fields is another. I have worked on projects that did need to store objects with different fields together but it seems to be a rare case.
From what I've seen, most projects actually have pretty standard objects with specific fields. People just don't like feeling constrained but in the end I don't think managing a schema is difficult at all. I actually like the safety and performance you get when you put in the effort.
I think the single most useful thing to see when starting a new dev job somewhere are the database schemas. If done properly, it will essentially tell you the story of that company and their data. You can think up questions and get answers such as "Is it possible for a Customer to have multiple Orders? Can a product have more than one Review?" (Stupid examples but they illustrate the kinds of relationships which are asked about all the time).
I guess my point is that the SQL world is still innovating but it doesn't get that much attention. Unless I have a specific use-case for storing similar entities together with different fields then a schemaless DB is not the right tool for the job.
Instead, you're writing a bunch of code to deal with data that may be in the old format, or may be in the new format, or may be in the new-new format.
I don't see that as an improvement.