Why MongoDB is a bad choice for storing our scraped data
blog.scrapinghub.com
blog.scrapinghub.com
It seems like it's a elegant technological metaphor (lets use mmap, the OS is our cache and we can overwrite in place in RAM) that in practise turns out to be a terrible idea. Overwrite/mmap cannot be made reliable, requires blocking write-locks, wastes disk, and causes problems shuffling data around as it grows. Add other bad decisions (keys aren't interned, seriously?)and it's just a terrible limping monster.
Abandon it, walk away.
MongoDB is every developer's wet dream. With it's expressive query syntax and extreme ease of use, everyone wants to drink the koolaid. This is a huge problem, because mongodb as a database is dangerous
[..]
I have developers begging me to let them use it. This time to collect logs from our servers for analysis later. I cave in, and give my go ahead, with a warning saying that no critical data can enter that section. Mongo processes were crashing. Several times per day. About 20% of the crashes yielded a completely corrupted database. This programmers wet dream quickly shows itself to be a serious operations nightmare.
http://hackingdistributed.com/2013/02/07/10gen-response/#com...
I have used mongodb for a number of smaller projects, and I have had an excellent experience. It's not "a terrible idea in practice". It might be a terrible fit for what you want, but that doesn't mean it's bad technology.
When you're not operating at significant scale, you can use a relational database. They're easy and fast to set up, have nice write-safety guarantees, are more flexible than a key-value store, and will scale well beyond anything that mongo has ever achieved. You can even use them as a key-value store! The downside, of course, is that you have to a tiny bit of knowledge about set theory, and that's a deal breaker for most "developers" today.
The whole point of the GP was that Mongo isn't elegant or easy...it's just naive and short-sighted, and the architectural mistakes within it are fundamental and probably unfixable (at least, not without killing the speed advantages they claim). The real reason that people use mongo is that most webapp devs don't have a very good understanding of how computers work, and want everything to look like Javascript, because that's all they really know.
Both these applications have been in production for about a year without a single problem. Both are using the same MongoDB instance - data size is about 200 GB (RAM on the Mongo machine is 16 GB)
Both these applications were previously on Oracle and were a pain to maintain. The Mongo schema is simpler and far more maintainable then the RDBMS schema. Backups/Monitoring/Replication on Mongo has never given us any problems.
Now, since you claim that an RDBMS can do anything better than MongoDB, can you point me to a simple/elegant/maintainable RDBMS schema for an e-Commerce Product Catalog? I would love to see one.
The original post, like most of the 'Why we moved away from MongoDB' posts displays a shocking lack of due-diligence on the part of the development team / tech lead at these firms. All the points under the 'Data that should be good, ends up bad!' section are known facts about MongoDB. All of them are covered in the manual. If you are not fine with any of these points - please don't use MongoDB at all. Don't put it in production. It baffles me how these firms can put MongoDB into production and 'discover' these things later. Instead of ranting at MongoDB, the CTO's of all these firms deserve the sack for lack of due-diligence and putting data at risk.
One last point:
> ..more flexible than a key-value store
MongoDB is not a key-value store.
Edit: And this is downvoted for calling out the fact that people on HN can't discuss a freakin' database without hurling insults.
The primary problem in software today is that we've confused the ability to build something with actually knowing anything of value.
Am I reading this correctly? It seems to imply that the ability to build something is somehow orthogonal to knowledge of value.
I don't know about throwing mud into a heap and then calling it sculpture, but if we are talking about the subset of "things" that have value in and of themselves, the ability to build them does imply some knowledge of value.
Now, the relative value of knowing how to put together simple web sites using jQuery vs. the knowledge to discover Chaitin's Constant is very much worthy of discussion. But likewise, the knowledge of how to construct a true but unproveable statement in a toy system vs the knowledge of how to build VisiCalc and revolutionize programming is worthy of discussion as well.
Not only are you reading it correctly, that is in fact (part of) what I'm saying. Building something doesn't automatically create value. We've confused the two.
I'll agree MongoDB is not a terribly well engineered database however I don't agree that SQL is always the best alternative in these simpler scenarios. There are lots of things wrong with using SQL to solve every data problem. I also don't agree that knowing SQL is equivalent to understanding "set theory". I've known plenty of DBAs who don't know the first thing about set theory and really just know just enough to install the software and piece together the APIs. The fact that one has chosen SQL doesn't make them good at working with data any more than choosing MongoDB makes someone bad, or implies they don't understand "set theory".
What concerns me mostly though isn't the MongoDB issue but that we can't discus the issue in a professional way, without elitism and distain dripping through. You seem to have confused "computer science" with "anything of value". I happen to believe there are more things worth knowing that are of value to practical software development than just the computer science (not to minimize that of course).
It's not about elitism, it's about making good decisions. That said, given all the hype with companies that hire, protesting loudly might not be the best short term personal decision. Meh.
And of course it is about elitism. Listen to yourself. "It's about making good decisions".
I mean who are you to judge from the outside what technology a company should use for a specific use cases ?
In the real world, people have been using relational databases to solve problems for years. They work, they're understood, they scale.
"global write-locking will not _inevitably_ lead to consistency or throughput unless the write frequencies are sufficient to cause and require that"
In which case, you can just as easily use a relational database and avoid the chance of problems altogether.
And they're a pain the ass and don't mix well with the kinds of programs many want to write. Mongo clearly fills a niche that relational databases don't serve well; if it didn't, no one would use it.
I still think MongoDB is great for many applications as many companies are using it for their data needs (like Foresquare), and the same with RDBMS like MySQL, that lot of big fishes use them for different parts of their architecture (facebook, twitter, etc). In the end, each option has pros/cons, but one will be better for your use case.
The idea that "it must be good for something or people wouldn't use it" is absurd. People do the wrong thing all the time. People make technical decisions based on fads constantly. Mongodb is one of the prime examples of fad driven development choices, where people choose it because "it is web scale" while having no idea what they are even supposed to be comparing it to.
You really think tossing out insults like that is a way to have a reasoned conversation? I think not, come back when you can converse like an adult.
> And I have no idea what "don't mix well with the kinds of programs.." is supposed to mean.
Then you need more experience as a programmer perhaps. I'm a programer and a SQL guy, and I'm fully aware of what a pain SQL can be in an application and if you don't see the ease of programming things like NoSQL databases or Mongo bring to programming, you aren't paying attention or you're lying to yourself about how well SQL fits with code.
In short you are both right, applications and also use of said applications and requirments make the difference as to what interface is best. Don't need to argue over that without even stating it. You both win, this is the internet - now laugh :).
The fact that you have some unspecified problem does not mean anyone else who does not have that problem is inexperienced. Given the complete lack of information available, it is just as reasonable to conclude that you are in fact lacking in experience which allows others to solve the problem you continue to refuse to define.
I say payroll databases at dawn, 50 paces each, turn and shoot. Go :-)
Both of them. You only wrote two sentences, it shouldn't be hard to find them.
>the one that many people find
You didn't say anything about "many people find". You said they are a pain in the ass, and don't mix well with the applications many people are writing. Those are both assertions, and you supported neither of them. Even after replying twice, you still haven't even given a hint as to what you might be referring to. That really makes it seem like you are just saying things out of ignorance.
I said they were a pain in the ass for the kinds of programs many people want to write. I'm sorry you're too ignorant to grok my meaning without it being explained to you like a five year old child.
> Those are both assertions, and you supported neither of them.
They are self evident facts and don't require supporting evidence; the very fact that a community exists around these products should make that clear to you.
In any case, it's absolutely clear there's no value in conversing with you, good day.
And refuse to specify what those kinds of programs might be. You are inventing a "many" and giving them a problem to create a false impression of consensus, when it is actually just you making a singular, baseless assertion.
>the very fact that a community exists around these products should make that clear to you.
I address that in the first post. You keep responding purely to act like a petulant child, but provide absolutely nothing to support your claims. Do you really think that makes you appear to be the rational, logical party?
Some compare it to relational databases... http://www.mongodb-is-web-scale.com/
Those sorts of people tend to be very focussed on clever algorithms and data structures, and frequently miss the larger picture and coding best practices. I've seen far too much code that had extensive CS cleverness at the root but was spaghettified, untested, undocumented, poorly performant, and not even tracked in a version control system. Such people often don't value writing code that other developers can read and maintain. On a team, clarity matters more than cleverness.
Good developers need both sets of skills.
Saying that I think it is a bit elitist to think you need a new term, who gets to say who can use that term?
Write locking definitely leads to throughput issues but it results in better consistency not less.
The original comment you're replying to aside, it was likely because stating that you dismissed the entire comment without giving an actual objection to it added nothing at all to the discussion.
Yes, if you know how to do it. But posts like https://news.ycombinator.com/item?id=5675902 and questions 'Should I learn SQL?' here and there make me think that's not required knowledge these days.
I still say it's bad technology. Use plain old SQL instead.
create table keyvalue (key text, value text);
create index keyvalue_idx on keyvalue(key);
And use is trivial as well: insert into keyvalue (key, value) values ('key', 'value');
select value from keyvalue where key = 'key';
[edit] formatting fix.They are also on my blacklist, forever.
They used to ship with unacknowledged writes as the default option. Think about it for a little, a database that just throws your _data_ over the fence and prays for the best, proceeding without an write acknowledgement.
There was not flashing warning on their front page about, no bold disclaimers, but there were sure plenty of "Oh look super fast benchmarks beating SQL and other NoSQL database, albeit created by fanboys".
That decision told me the story of who they are and what kind of principles they use to build their product. It wasn't a mistake it was a deliberate shady tactic employed.
(Yes, I know I have written about it at least 3 times before and will mention it every time I see MongoDB mentioned )
books: [
{id: 1, tags: ['a', 'b'], author: ['c', 'd'], count_read: 123, count_bought: 456},
]
where count_read and count_bought is atomic increment, tags are arbitrary string array and could be indexed for searching.Yeah I knew Postgre could do that. But MongoDB is the most simple and direct way on the market. Stuff like tags makes MySQL m2m joins very inefficient.
Really? When I don't want to think in databases and only on persisting my native types in my favorite programming language MongoDB is my choice. I use it only for experimenting, so I don't care about all those scalability issues.
Was your comment a little bit ironic?
s/MongoDB/Any technology/g
The size of the niche varies. The lesson is to be sure the choices you make are appropriate for your situation, and be aware that things may change if new requirements emerge or scale needs to go beyond what you projected. These concerns are not specific to MongoDB. It is a rare project that goes from prototype to small scale to large scale on its original implementation technology choices.
All easy things get us in troubles when we want high performance. Cheers;)
MongoDB makes it easy to scale out (replica sets and sharding), where is the "easy to setup replicated and sharded open source SQL database?"
I mean, I know that Postgres has replication (via Slony? honestly, it's been awhile since I looked at their solutions) but I don't recall it being as dead simple to set up.
For me, setting up replication needs to be easy because we redistribute the store as part of our product and we need scalability (both replication for redundancy and sharding for scaling).
So I'm honestly asking here, where is the easy to use sharded and replicated open source SQL store that I've been missing?
For example, today my entire production MongoDB database was running 3x slower because a single replica in one shard was down, and their buggy PHP driver kept trying to talk to it despite it being marked down. I really enjoyed waking up at 2am to deal with that.
It relates back to the "easy to use" nature of their marketing. It really is super easy to use and develop on, but the minute you need to do anything important or serious, it breaks down.
You aren't doing yourself any favors going with it except as a proof-of-concept.
FWIW, it's been spotless for us so far. Our needs aren't web scale, but they're big enough to need scaling features.
So, it's less "nosql vs. sql" and more just "don't use mongodb".
Sharding certainly isn't as easy - the technical compromises that mongo makes make it pretty trivial to implement, whereas it's relatively hard to make it work in an RDBMS while maintaining all the expected capabilities. It generally requires some application-level work on open source dbs.
With that said, I really think many people grossly underestimate the effectiveness of scale-up. It's worth remembering that Stack Overflow (for example) is still running on a single pair of master/hot standby database machines.
And I think you underestimate the benefits of scaling out. If I want to ensure close to 100% uptime or have a server closer to my users than Cassandra or even MongoDB would be infinitely easier to setup and manage than Postgres. These are "very nice to haves" for even the tiniest startup.
You can replicate much of this behaviour using open source RDBMSs, but yeah, it's not what they're designed for and it's harder. If you want quality replication/clustering you're currently looking at paid-for DBs.
Being able to scale out is absolutely a nice-to-have. I'm not sure it's a nice-to-have on the scale of giving up all of the features an RDBMS provides for most people's use-cases. Further, you might find that the relative lack of data headaches you get with an RDBMS more than makes up for a little extra time setting up hot standby.
Finally, if 100% uptime is that important you're probably not relying on a relatively niche NoSQL database. If uptime on the level of Stack Overflow is good enough (which for most people it probably is, let's face it), then you'll probably find replicated postgres good enough.
When I was looking at implementing Cassandra instead of Mongo DB, it seemed like we had to create reverse column family (IIRC, been away from Cassandra for a bit now). Is that still the case?
Has the state of affairs advanced since a year ago? Would love to hear it has!
[1] http://brianoneill.blogspot.com/2012/03/cassandra-indexing-g...
- Ordered data and skip / limit: These would run just fine on any database system. Given that you have appropriate indexes. It does not matter if there are a trillion items total, as long as you are seeking over an index and the result set is in reasonable size.
- Restrictions: A lot of software has restrictions. Filesystems has file name limitations. RDBMSs have table / column name limitations. It's a fact of life. Why is this a con for MongoDB?
- Impossible to keep working set in memory: It is a fair argument that MongoDB has shitty memory management because it just delegates the responsibility to OS. However, this is a concern with any DBMS. Also, given that there are appropriate indexes, you don't need to keep the entire database on memory. This comes back to indexing problem.
- No transactions / lack of schema / no joins...: I don't remember mongoDB claiming to have such features. My car can't fly. I'm not complaining. (Well, sometimes)
- Locking: Fair point. Better I/O performance might come handy (like an SSD) or eventually sharding.
- Poor space efficiency: Fair point about fragmentation and field names. Compression can be achieved on the filesystem level. There was an article about that a couple of days ago. I'm not sure about pefroamnce though.
- Too many databases: This should not be a big issue. Mongo does not go ahead and allocate a couple gigagbytes for each db, it uses incremental file sizes.
- Silent failures: Yep.. There it fails miserably. Recent versions are better though.
I just don't like people bashing something without valid reasons. It might just be a perfect solution for similar applications, this is not a good way to evaluate.
That really doesn't seem to be the case here. Like you, the article's author(and several others here) have had issues with it for their particular use-case, and the reasons are clearly listed in a well organized paragraph by paragraph summary explanation in the article. Others here who've had a similar experience at least stated they had issues with it as well, even if they didn't go into much detail about it.
And speaking of the lack of valid reasons, to be fair, many relatively new technologies like these often get significant praise/hype without many valid reasons as well, other than [X]startup/company is using it, so it should be able to work for me, or it must be an awesome technology to use.
FWIW, we still use Mongo in other internal applications, it's just not the right choice for our crawl data storage backend.
Transactions for example have never existed in MongoDB and joins doesn't really make much sense.
After all, developers are rather susceptible to the "don't tell me I can't do that" behavior.
At some point you must have compared it to, say, Postgres – which is what the section before the summary hints to.
All the HN crowd went insane at once or, who knows, maybe they found out one product that is relatively heavily marketed is mostly blowing smoke up everyone's asses.
The complaining is vis-a-vis the marketing and perceived fan-boyism. It might also not be completely bad news as it means people are still using it.
""" Ordered data
Some data (e.g. crawl logs) needs to be returned in the order it was written. Retrieving data in order requires sorting which is impractical when the number of records gets large. ""
it requires _indexing_ and is quite feasable as I do it every day with stock ticker logs ( also required to be retrieved incrementially )
There are a few other flags that make me wonder about the exact limitations you found, but I will be anticipating your follow up post to see what your fix was since some of those issues are very common.
He mentions the lack of joins, but doesn't say a word about Mapreduce.
"MongoDB needs to walk the index from the beginning to the offset..." You don't "walk an index". It's an index.
"Too many databases" sounds a little suspicious. Why not add an indexed field to partition records?
Complaining about a lack of schema, transactions and triggers? Really? Did you read the docs at all before starting?
MongoDB is not without its problems, but friend, I think you wanted either Postgres or Hadoop.
If you have an address book, you don't have to walk through the city to find an address, but you do have to look through your address book in some way or other. Of course, you can have an index of the index ("C starts at page 7"), but then you have to look through the index of the index.
HBase is multidimensional though, which allows you to keep N numbers of versions of a cell. By default you will get the latest version of the cell back, but you could also opt to receive N versions back, which is useful for time series use cases.
As far as I know, key design is not an important aspect with MongoDB but I could be mistaken. HBase has a pretty awesome book (http://www.hbasebook.com/), which has an entire chapter dedicated to key design. Lars (the author) also has a pretty in depth 1 hour video on key design (http://www.youtube.com/watch?v=_HLoH_PgrLk).
HBase is pretty widely used, I've seen 1200+ node clusters running production tables.
It's a good point that some of this can be achieved with indexing, I should have given more details in the blog post.
What interests me is why they would want to keep everything in the database? I'd assume that they need to aggregate and curate the scraped data. After the initial scrape the majority of actions surely are going to be on the metadata of the scraped content? (where is said data, when was it scraped, how big, relationship to other data, etc) This data is much smaller and can be stored in relational database, as its proper structured data with relationships.
This allows the nasty unstructured data to be kept on a plain boring filesystem. After all filesystems are exceptionally mature, universal, multilevel key-value stores.
Now people will say that filesystems don't scale, well that's not really true. ext4/ntfs on a single system won't scale, but something like lustre/gluster(although not as neat)/gpfs scales linearly with the amount of nodes you apply to it.
We'd need to code the searching, filtering, paginating, (distributed?) job management ourselves while being careful to keep the DB & metadata consistent. It works best if each file is a reasonable 'chunk' of data (not too big, not tiny). None of this is a problem, and it scales very well as you said.
In the end, we went with HBase for crawl data in the new system. Of course, you can look at this as files on a filesystem (HDFS or others) :) It does a lot of what we would otherwise have to code ourselves and it's a good fit for applications we want to build on that data in future (e.g. storing other crawl datastructures, processing with hadoop). I'll provide more details on that in the next post.
"The lack of joins & transactions of course did factor into the original decision. My point (which perhaps could be clearer) was that MongoDB ended up being used outside of the area in which we originally intended to use it. There was some reluctance to add another technology when we could get by with what we had for what was (initially) only a small use. Additionally, some limitations were not always well understood by web developers (who were new to mongo and enthusiastic to try it). I see this as our mistake. With hindsight, it’s clear we should have introduced an RDBMS immediately and kept MongoDB for managing the crawl data."
So I think it's reasonable to be cautious about using a system that can't effectively grow outside of its initial special purpose.
The real issue here is that it feels like the author has just 'discovered' these problems as if Mongo was hiding them all along and after a long time using the system he just found them. The reality is that all of the things he brings up are well documented. It is fascinating to me how people pick a buzzword database and don't bother to think about how their application might run poorly on it over time.
Firstly, the fact that some drawback is well documented does not excuse the fact that it is a drawback.
Second, while some drawbacks are documented some implications of these drawbacks are nuanced and only become obvious with experience. A good example of this is the implications of "schemaless" databases (more accurately: databases that do not check data against a schema). Not having to migrate tables is a boon for lots of development. It's also a giant pain if it turns out that bugs cause data integrity issues.
Third, this experience report is really useful since poorly structured scrape data is one of the areas that I would have considered to be ideal for mongodb.
Most people don't have perfect foresight. I don't fault the author on his lack of omniscience with respect to how mongodb would turn out for them. His original reasoning (given in paragraph 1, sentence 1) does not seem stupid.
https://jira.mongodb.org/browse/SERVER-2986?page=com.atlassi...
Did you guys roll your own HBase environment or did you go with the CDH? If you're using the CDH version and have any questions, feel free to shoot an email to cdh-user.
Cloudera has in fact been an inspiration for us to follow, you guys have really struck the right balance between open source and commercial support. We follow the same philosophy with Scrapy (an open source web crawling framework), as you do with Hadoop and its ecosystem.
I'm not too familiar with Cassandra, but the scalability of an HBase table is almost entirely dependent on your key design. Judging from their use case and requirements, they would likely use a incremental key design which would allow for super fast range scans, of course, this leads to region server hotspotting, which may or not may not be a big deal to them.