Guide to MongoDB for startups
optinidus.com
optinidus.com
* Do you have a customer yet? If not, technology choices do not matter, go build your product as fast as possible and get a customer.
* Is your system starting to have slow performance during max usage? If so, every system will have a few easy optimizations, find those using something like NewRelic. Aside from those, technology choices do not matter, throw money at any scaling problem (at this point, it typically isn't that much money). This will optimize your time for sales and marketing to get more customers.
* Is your system starting to buckle under the weight of customers? If so, go hire someone who can scale everything you learned from your customers, and not someone who bends to the latest trends on HN.
Here is an article on my philosophy and experience with scaling businesses and architectures: http://blog.mongohq.com/changing-the-growth-formula/
We use MongoDB in production, processing millions of new records every month and it works great. Our data has differing and evolving schemas and can be used as JavaScript objects or python dictionaries with little effort. It works and performs well for us and our use case.
I'd be delighted to find out what these nay-sayers are doing that MongoDB must fail so incredibly at for them to have such strong negative opinions.
I'm not speaking only to the existing comments on this post but the general theme of past MongoDB submissions.
Personally, my 10k foot impression has been that MongoDB is really easy to get started with, has some cool features, and is fun to use... but then in production it has serious issues that are fundamentally unworkable. Until I see a series of: "Here is what was wrong with MongoDB and how we fixed it" articles floating around along with a number of endorsements - I'm probably going to keep that 10k foot view.
I could easily be criticized for not digging in to Mongo's problems myself to learn exactly what they are and whether or not they are still an issue, but life is too short and there are too many other things to do.
100% agree. I think it's become cool to knock MongoDB, just because it is MongoDB.
It isn't a silver bullet, and its use cases are more limited than RDBMs. But if you know what you're doing, it has some great applications.
You really don't think it has anything to do with the numerous serious technical deficiencies it suffers from? (Read some of the other comments for this submission if you really do need more details about these problems.)
When it comes to databases, fooling around just isn't an option. If a given database technology has flaws, then those flaws need to be widely known. And this isn't done to be "cool". This is done because the safety and integrity of data are of the utmost importance when working with database systems, and it's absolutely critical to know about anything that may cause problems.
Yes, I do, actually. I see a TON of out-of-date commentary on places like HN, or just stuff that is factually incorrectly.
I have used MongoDB in production for 3+ years, and it hasn't caused any problems.
I will say the out-of-the-box set up isn't what you want in production.
Just because your small scale system is running okay doesn't mean MongoDB is great, just that it's okay for your application.
The issues encountered ranged from MongoDB failing to scale out as easily as promised, to significant loss of data. There was also some backlash because MongoDB didn't (and perhaps still doesn't) persist data to disk when it acknowledged it as received[0].
A general perception grew, rightly or not, than MongoDB was being marketed by 10Gen[1] over and above its capabilities. Then the actual message of "MongoDB is not a drop-in replacement for traditional RDBMSs like PostgreSQL and MySQL, and is not the ideal solution to every problem" has been filtered down to "MongoDB is a terrible product with no use cases whatsoever." Such is the effect of the HN echo chamber.
Much like PHP, MongoDB is now simply a product that you cannot say anything good about on HN, even if you are in fact finding it an effective tool for your use case in spite of its shortcomings.
[0] Cynically, because they were trying to win benchmarks. Less cynically, because the use case was semi-ephemeral data where possibly losing some data is acceptable.
[1] Now also called MongoDB
This is probably what raised awareness of the issues of MongoDB http://www.slideshare.net/emiltamas/scaling-with-mongo-db-wi... and it's from 2011
It's a lot of time for a software project
Once a DBMS like MongoDB gets popular, then people uncover some of its failings and write about them...which is great because it means you don't have to discover the same things the hard way. But it's a drag too because there's no database system that everyone agrees is adequate (much less ideal) for every application.
There's a neat helper library in nano, but that just constructs requests for you.
I will grant that "don't use NoSQL" isn't the correct advice in absolutely all cases. Once in a blue moon there is a case where it's the best solution, and if you say you have one of those cases, fair enough, I'll take your word for it.
But the vast majority of people who use NoSQL should be using a relational database instead. And the consequences of using NoSQL where you shouldn't, can be very bad indeed. (If your data gets messed up, it's not like a bug that can be put right just by fixing your code.)
So I will stand by "don't use NoSQL" as general advice. If you're really sure you know what you're doing, if you have sufficient expertise on databases and understanding of the present and future characteristics of your workload to be certain you're an exception, fine; but then you don't need general advice in the first place.
"Firebase stores all its data in MongoDB which offers the built-in capability for apps to scale automatically and gives each piece of data its own unique URL stored in JSON documents."
I was going to see if I can get achieve data consistency problems there for my own firebase app.
Also posts that malign MongoDB frequently make it to the home page [3][4].
[1] http://aphyr.com/posts/284-call-me-maybe-mongodb
[2] http://www.mongodb.org/about/production-deployments/
[3] http://www.sarahmei.com/blog/2013/11/11/why-you-should-never...
[4] http://www.aaronstannard.com/post/2013/12/19/The-Taxonomy-of...
For a group of people that know nothing but hammers, screwdrivers don't make much sense. This feels often what these "bash Mongo" arguments sound like.
There is no silver bullet. Just flippin' understand the tool you are using and how to use it.
If people understand Mongo and its shortcomings, then fair enough. The thing is, those shortcomings are large and far-reaching, and the number of applications for which Mongo is a truly superior long-term choice are relatively few. It's my perception that when the hype around Mongo peaked there was an awful lot of ignorance surrounding the product - the current behaviour towards it is probably a reaction to that.
So, what is your stake in it? You like it? You have a lot of customers on it and you want to validate your decision?
I'm just curious, because by my tastes, MongoDB is an awful database. It's like the PHP of Databases. It seems like there are a lot of better choices out there especially in the last few years.
(Also, some guy by the name of Zuckerberg made a TON of money with PHP.)
Also, the tool is still in its toddler years ... and it is constantly compared to tools that have been around years longer (PG/MySQL). Comparing it to other NoSQL databases is just as silly (Riak, Cassandra) because they are completely different, have different constraints and are solving different problems.
There aren't better choices ... there are just different tools that are better for different jobs. It would be like looking at your toolbox and chunking everything but your saw or your hammer. Too many people are looking for the "technology to rule them all", but that doesn't (and shouldn't) exist.
It fails miserably in various ways when processing millions of new records per hour. On a very large scale MongoDB is unusable, but it is fine for problems 2-3 orders of magnitude smaller, so long as the data is not mission-critical.
It is indeed the millions-per-hour order of magnitude where things get interesting.
They'd be shocked(!) at the news. ;)
When an article opens with a statement like it's difficult to give it much credence at that point. Unfortunately, this ignorance is very common among the MongoDB crowd.
Like your example? millions of new records every month? That's not a "huge amount of data". We're adding tens of millions new records each day, into mysql, and I would still consider our usage extremely intermediate if not amateur.
MongoDB has many benefits, they've improved on numerous oversights and flaws in their initial design, and it's becoming better all the time. What I'd love to actually see change is the naive belief of MongoDB users that it's a new, unique or vastly superior solution.
Since I'm taking a few hits on "millions" I should perhaps clarify it's more like "hundreds of millions". Still obviously not a gigantic amount but certainly enough for me to be able to say that it is functional enough and robust enough as a database that a lot of the negativity is if not unwarranted, then at least excessive.
There's just not much to say tho'. If this were an aviation site and people kept posting about how piston engines were the future because jets "don't scale to large aircraft" and "were never designed for long journeys" how would you argue against that? There would just be no point.
I'll say this tho': in 1969 IBM released a NoSQL product called IMS, which was designed to handle the BOM for the Saturn project. IMS has been continuously developed since then. Have a look at it and what it can do, then look again at MongoDB, which is - genuinely - about 40 years behind the state-of-the-art. Then understand why experienced database guys* can't be bothered to debunk it any more.
* my own stats: 50+Tb in Oracle with 800G changes/day, replicated to 2 other identical databases over WAN links. 2+Tb in MySQL. 10,000 commits/sec + 50,000 reads/sec on Oracle. 99.99% uptime on all of this. When people tell me "SQL databases" don't scale or aren't reliable, I laugh. And then ignore them.
Also, the various gotchas involved with maintaining MongoDB at scale meant I spent most of my team's time chasing down esoteric problems instead of serving customers.
But ultimately, trusting a database with decades of development behind it to handle something like a JOIN turns out to be a good idea.
In my experience.
Don't.
Or did you instead have a detailed and carefully plotted comment about the problems with Mongodb, but then you read "nasalgoat's" comment and found his concision so bold and bracing that you were intimidated from contributing your comment?
If so, I urge you to reconsider.
https://blog.serverdensity.com/tech-behind-time-series-graph...
This is nothing wrong with MongoDB.
1. Don’t do it
2. [for experts only] Don’t do it yet
What do you (or others) suggest as a good replacement for MongoDB in that stack?
Picking something because it's easy is bad.
I'm actually a bit surprised how long it took me to figure out that's actually the best practice -- I dug around a bunch of rails-like frameworks and was wondering why they didn't seem to be more widely adopted (Sails, Compound, etc) until it finally made sense that "The Node Way" is to add bits and pieces (usually on top of Express) to build a stack that's specific to needs.
What? There is no limit how much data you can stuff into a relational database and it mostly works really good. There are definitely some scenarios where relational databases are not optimal, for example aggregating huge datasets, but there are often solutions like materialized views or aggregates. The statement that RDBMSs are not designed to handle huge datasets is plain wrong.
With these growing needs the it was getting even more difficult to define a fixed structure to the data and the need to have a solution which handles unstructured data grew even more.
I don't understand how once inability to come up with a good data model is related. If you have no data model you will have a hard time working with your data anyway. You can stuff unstructured data into a string column and modern RDBMSs also support semistructured data like XML and JSON including operations to manipulate such data. And even with XML and JSON you still have a data model, just not a relational one.
I agree that the relational model is somewhat stiff and it is - somewhere between sometimes and often - a pain to design, implement or evolve a schema. But a fair amount of the complexity usually comes from the domain you are modeling and does not go away when you switch to a different data model - your database may not complain when you stuff documents with seven different schemas into it, but now the burden is on the code to deal with documents with seven different formats. Data modeling and model evolution is a difficult problem and does not go away by switching technologies. The major difference is the point in time when you recognize that you have a problem.
So most of the NoSql databases traded consistency for Availability and Partition Tolerance which was a fundamental shift from the relational world where in you had to design your schema perfectly so that there were no inconsistencies in the data ( 1NF, 2NF, 3NF, BCNF).
Consistency in the sense of the CAP theorem and in the sense of database schema normalization are almost completely unrelated. This gives me the impression that the author does not really know what he is talking about.
One of the biggest advantages claimed by NoSql databases is Horizontal scalability , which is the ability to handle more load by adding in more machines, since the computing power these days has gone cheap compared to the effort required in to fine tune the app and add more computation power and memory to an existing machine.
RDBMSs support partitioning across server as well. It is probably harder to set up and does not scale as well because the provided guarantees are stronger, but it is possible.
MongoDB uses reader-writer locking mechanism , it gives concurrent access to reads but exclusive rights for write operations which means it can handle concurrent read operations but if there is aright operation it will block the reads until it gets completed .
So do many RDBMSs. Or they use MVCC and allow even more parallelism.
Prior to mongodb version 2.2 mongodb had an instance level lock which means that whenever there was a write operation it used to lock the entire mongodb instance and even if there was a read queued for a different database it will have to wait as the write operation blocked the entire mongod instance.
This was changed in 2.2 where in the write operation locked the whole database instead of the complete instance. So the solution which was left to scale the write operations was to add in more shards and route the next write query to a separate mongodb shard instance.
This is still extremely inferior to modern RDBMSs which usually support row level looking.
For applications this is by far the most important point to take into consideration when choosing which database to use, as for write intensive application this might come in their way to scale.
RDBMSs allow you to opt out of consistency - if you want to read uncommited data, you are usually free to do so. I don't see why a RDBMSs should intrinsically allow less parallelism but admittedly dropping consistency guarantees is quite contrary to the reasons you usually choose a RDBMSs to begin with.
UPDATE: I probably misinterpreter the last part about locking and scalability - after reading it again, it sounds like the author actually warns that using MongoDB may cause scalability issues. I leave my comments as they are although they sound a bit strange in that light.
TO-THE-DOWN-VOTERS: I would love to hear where I am wrong, my opinion is neither set in stone nor absolute truth.
> The statement that RDBMs are not designed to handle huge datasets is plain wrong.
It's especially strange given Mongo's less-than-stellar scaling.
I suppose row-level locking is a necessity when you want to be able to scale vertically. By comparison, I understand that MongoDB doesn't even try to support vertical scaling so it makes sense to not bother with complex locking systems.
Personally, I'd prefer to have to option of scaling without being forced to increase the number of moving parts in my single most important system (the database)...but the MongoDB docs make the fair point that there's a ceiling to vertical scaling. It's probably higher than most people think though.
As an aside, I discovered recently that Informix supports byte-level locks for 'smart large objects' (namely CLOBs and BLOBs). Makes me wonder whether field level locking will appear some day.
Locking has a degree of complexity to it, but implementing page locking (for example) is not that hard. The fact that the Mongo guys haven't is probably symptomatic of the fact that they're still relying on the OS to cache data, rather than implementing their own page manager like most other DBMSs. That said, given the short duration of Mongo locks, lock granularity isn't as big a deal with Mongo as some make it out to be.
I worked at Wine Spectator for a year (2010-2011), http://www.winespectator.com/ . They had built their first web site circa 2000 using Oracle, Sun Solaris, and Vignette with Java templates. Circa 2009 they decided to scrap the old, expensive system and move to PHP/MySql and the Symfony framework. They could not decide what their new schema should be, so, in the name of keeping things flexible, they decided that all data would be in a single table. This table had 240 fields, most of them with generic names such as "modifier_01" and "modifier_02". This was an organization that was looking for the flexibility offered by MongoDB, but they tried to cram that flexibility into a relational database, and they did so by ignoring all the relational features offered by MySql. This "one size fits all" database table did not work for the last assignment I was given: import all the old FileMaker Pro databases to a system running Mysql/PHP/Symfony. I build an entirely different project, with its own database and schema. Lord knows who is maintaining it now.
Then I worked a year (2011-2012) at Shermans Travel, http://www.shermanstravel.com/ , which also tore apart its database. When I first arrived they were trying to save a system built with MySql and PHP and later forced to conform to the CakePHP framework. The database had over 300 tables, many of which were no longer in use. Of the tables that were in use, many had fields that were no longer in use. The code, and the database, were a sprawling mess, that had evolved chaotically. (I am emphasizing this chaos because this criticism is often made of MongoDB: without a schema then how do you keep your data organized? Well, most of the places I have worked have had relational databases where the data was completely disorganized). After a few months, the CTO and the tech team decided on a complete re-write of the code. The tech team was allowed to vote for either Java or Python or Ruby (no one wanted to use PHP). We voted for Ruby. We rebuilt the site as 6 apps, using MySql for some of the apps and MongoDB for some. My last big project there was a rescue effort for a broken group of 4 database tables in MySql. There was a "users_history" table that was suppose to track whether a user had subscribed or unsubscribed to various newsletters we offered, but there had been a bug, apparently for years, such that many of the "unsubscribe" attempts were not recorded. There were 3 other tables with somewhat redundant data, and I wrote a script that scanned those other 3 tables and attempted to funnel the correct data to the 4th table.
I have many more stories like this. I could write a whole book about places where Oracle, MySql or PostGre was in use, but the data was badly organized. Unused tables and unused fields are extremely common.
Why do I emphasize the chaos I have encountered? Because the charge of badly organized data gets thrown at MongoDB a lot. If you would like to read a scathing attack against MongoDB, read this:
http://www.sarahmei.com/blog/2013/11/11/why-you-should-never...
But to me, this line of argument compares the platonic ideal of relational data against the actual use of MongoDB. Maybe if Edgar F. Codd designed your schema to the 7th Normal Form then your schema really is well organized, but I have not seen anything like this in real life.
What I have seen, in real life, convinces me that every organization has an informal schema that is constantly evolving, and which can not be maintained with anything like regularity. Most of the organizations I've been hired at default to rebuilding everything every 5 or 6 years, because by that point the old system has grown chaotic. Sarah Mei's description of the dangers of MongoDB matches my own experience with relational databases: "we figured out that we had accidentally chosen a cache for our database."
What I like about MongoDB is it openly, boldly declares that the chaos I've seen is typical, and it facilitates the evolution of the schema which is going to happen no matter what you do or say. Things evolve, often chaotically. A programmer has a great idea and works on it in 2007, another programmer takes over in 2008, the project starts as raw PHP, later is imported into Symfony, then is re-written in Ruby, then it is broken up into several small apps. Someone quits. The CTO is fired. Someone new starts working and, for the sake of simplicity, prefers doing as much as possible as a background task, using cron scripts. A year later someone joins and is disgusted with the profusion of cron scripts, they want everything organized around a message queue. A very good sysadmin joins the team at a time when most of the programmers are weak, the sysadmin re-writes many of the background scripts, but he prefers Perl for everything and he implements some data caching strategies that no one understands.
You may think that I am exaggerating the level of chaos I have seen. I have not worked for Facebook or Google or Apple and if you tell me that in those companies everything is well run and well organized, then I will believe you, as I have no reason to doubt you. But I have worked at a lot of older media companies in New York City, and what I have seen is constant churn, churn at every level, churn in the team, churn in the technologies, and churn in the database.
I know I will be misunderstood, so let me try to clarify this:
I am not saying that chaos is good.
I am not saying that MongoDB is good because it encourages chaos.
I am saying that chaos is a symptom of the fact that most businesses do not know what their schema should be, and even if they did know what their schema should be, their needs would be different a year from now. The real schema needs of the organization (that is, what sets of data should be acquired and what the relations should be between those sets) are undergoing constant evolution, and this evolution is necessary, healthy, and unstoppable.
What is the strength of a relational system? Consider Wikipedia's explanation of Codd's Theorem:
http://en.wikipedia.org/wiki/Edgar_F._Codd
"The domain independent relational calculus queries are precisely those relational calculus queries that are invariant under choosing domains of values beyond those appearing in the database itself. That is, queries that may return different results for different domains are excluded. An example of such a forbidden query is the query "select all tuples other than those occurring in relation R", where R is a relation in the database. Assuming different domains, i.e., sets of atomic data items from which tuples can be constructed, this query returns different results and thus is clearly not domain independent."
Clearly, this assumes that the relations among the data are known. The organizations that I work with have no real idea about what relations they want to establish among their data. They are in a permanent exploratory phase. I believe these organizations could be described as "pre Codd", but most of them have been "pre Codd" for decades, and they will always be "pre Codd". If you force them to specify relations among their data, you will get answers exactly as useful as these:
http://blog.jimmyr.com/Funny_student_Exam_Answers_13_2008.ph...
MongoDB is useful in this context. Start acquiring data. Don't pretend you know what your schema is. You do not know what your schema is. The schema is changing all the time anyway.
Is there a place for relational databases? Yes, because sometimes some parts of the business become steady for some length of time, and for that part of the business, capturing fixed sets of data, with fixed relations, is very useful. But we should not pretend that this situation holds where it does not. I am not convinced that this is even the general case, though there is an overwhelming tendency in computer science, and in business, to pretend that fixed-sets-with-fixed-relations is the general case. If you feel it is, then you have been working at places facing conditions far more steady than what I have seen, or perhaps you are simply considering a shorter time frame than I am.
And without knowing your schema in advance, how do you ensure consistency? If you want consistency with mongo you basically need to either group all data that will get modified in any one logical operation into a single document, or you need to have (potentially extremely complex) strategies for resolving conflicts in between modifications to groups of documents. Achieving either implies a decent bit of domain knowledge.
I don't think it requires a whole lot more thought that you ought to be putting into your Mongo "schemata" anyway.
Now that kind of usage of relational databases will destroy MySQL since it doesn't do clever query planning.
But what is it that SQL databases could not solve which lead to the evolution of NoSql databases...{some "big data" changes to the world}...Which means processing huge amount of data. Which SQL databases were never designed for.
Firstly, specialized data storage and retrieval, including unstructured or document oriented, existed long before SQL did. Your filesystem is just such a system. Everything old is new again.
And secondly, a citation is required for the SQL databases "were never designed for" huge amounts of data claim. To start with, a definition of huge is necessary to make such a claim. 1GB? 1TB? 100TB? There are SQL solutions that work with all of those with ease. Many of the largest databases on the planet are humming away on SQL systems right now. SQL is abstract from the underlying platform, so if you have a cluster of 100 machines each in front of 1000 storage arrays each in front of 100 SSDs, it doesn't suddenly become "NoSQL".
SQL databases are generalized solutions. They generally do not solve specific, individual problems as well as precisely engineered solutions, which is why giant companies like Google, with extremely precise needs (e.g. index the web) have solutions that do a much better job for their purposes. Does that apply to your needs at all, though? Are you looking at your specific requirements, engineering precise storage and retrieval that is optimal? Probably not. Saying "MongoDB is under the umbrella of NoSQL, and NoSQL also kind of encompasses highly specialized solutions from industry leaders, so it will work for my startups contacts databases" is very, very poor, misleading reasoning. But we see it all of the time on HN in regards to solutions like mongodb. Which again is how you get responses that might seem hostile.