MongoDB Gotchas and How To Avoid Them
rsmith.co
rsmith.co
1. The keys in Mongo documents get repeated over and over again for every record in your collection (which makes sense when you remember that collections don't have a db-enforced schema). When you have millions of documents this really adds up. Consider adding an abstraction mapping layer of short 1-2 character keys to real keys in your business logic.
2. Mongo lies about being ready after an initial install. If you're trying to automate bringing mongo boxes up and down, you're going to run into the case where the mongo service says that it's ready, but in reality it's still preparing its preallocated journal. During this time, which can take up to 5-10 minutes based on your system specs, all of your connections will just hang and timeout. Either build the preallocated journal yourself and drop it in place before installing mongo, or touch the file locations if you don't mind the slight initial performance hit on that machine. (Note: not all installs will create a preallocated journal. Mongo tries to do a mini performance test on install to determine at runtime whether preallocating is better for your hardware or not. There's no way to force it one way or the other.)
#2 - usually the journal should be pretty quick to allocate - I've not experienced this problem directly myself.
I'll add some extra bits to the bottom of the post with your notes.
The Mongo team seem somewhat reluctant to implement it.
Compression could have drawbacks on documents that get updated frequently. But it will be extremely useful on documents that get created and rarely/never change, coincidentally what I mostly have.
It would also greatly help if keys are compressed or indexed in some way since it could be done transparently.
You may recall the Mongo team being reluctant to make the database function well in single server setups, but they did address that with journalling.
Also, I'm working on Snappy compression with Mongo (it's already used for the journal), however it's not currently stable and work is sporadic due to my startup.
In src/mongo/db/db.cpp I see:
void _initAndListen(int listenPort ) {
...
dur::startup(); // <- this i believe preallocs journal files
...
listen(listenPort); // after journal prealloc i think
p.s. this blog post is a very interesting article overall i think, without commenting on all the specifics...MongoDB does not support joins; If you need to retrieve data from more than one collection you must do more than one query ... you can generally redesign your schema ... you can de-normalize your data easily.
This is a much larger issue than it seems - nested collections aren't first class objects in MongoDB - the $ operator for querying into arrays only goes one level deep, amongst its other issues, meaning that often-times you must break things out into separate collections. This doesn't work either, though, as there are no cross-collection transactions, so if you need to break things into separate collections, you can't guarantee a write to each collection will go through properly. (Though, I suppose if you're using the latest version, you could lock your whole database)
Also, the postgres integration (linked/discussed) here on HN
1) Make sure to permanently increase the hard and soft limits for Linux open files and user processes for the MongoDB/Mongo user. If not, MongoDB will segfault under load and when that happens, the automatic recovery process works incredibly slowly. It's a bit tricky to get this right, depending on your level of sysadmin knowledge. 10gen doesn't emphasize or explain the issue very well in their docs: "Set file descriptor limit and user process limit to 4k+ (see etc/limits and ulimit)" That probably makes sense to just about 0.1% of the people setting up MongoDB: http://www.mongodb.org/display/DOCS/Production+Notes#Product...
2) Make sure to disable NUMA. This 10gen documentation note is a great example of clear documentation: "Linux, NUMA and MongoDB tend not to work well together ... Problems will manifest in strange ways, such as massive slow downs for periods of time or high system cpu time." Massive slowdowns and mysteriously pegged cpu usage on production database systems are definitely 'strange'. I would probably choose stronger and more precise language, but 10gen clearly knows what they're doing: http://www.mongodb.org/display/DOCS/NUMA
tl;dr If you have problems with MongoDB, you aren't using it right. Read the documentation more carefully, and then when that doesn't work, hire an expert.
I'm getting the idea that is rather challenging to use mongoDB right. While there's certainly a place for power tools that can only be used by highly trained experts or you risk disaster... that kind of goes against the idea that mongodb has anything to do with 'simplicity', don't it?
Nevertheless, a rundown of the gotchas and how to avoid them based on experience beyond simply running apt-get install mongodb is one of the most useful pieces on MongoDB I've seen of late.
The only new-news for me was that SSL support isn't compiled in by default. That's pretty irritating. I wonder if that applies just to 10gen's packages or also to distribution provided mongodb packages.
Edit: Link to applicable docs on how to compile in/use. http://docs.mongodb.org/manual/administration/ssl/
So it's available; albeit there is an intent to have a subscriber build with some extra features that are heavily enterprise biased in their usefulness.
Most of the distribution provided packages (well, ubuntu / debian) were horribly out of date last time I used them. Not sure if they do, or don't have SSL support - but I doubt it as most of them were < 1.8 which I think is pre-SSL (not 100% and can't find the commit though).
Oh, they still are. I use the 10gen apt source.
I thing this will be a good slogan - 'We are PHP of storage engines.'
Even if it's not your favorite technology, sometimes you end up in a position where the rest of the company is using something, and you need to work within those constraints. It's important to understand the technologies you're building on, their configuration options, and to understand the best practices way of working with them.
This, by the way, is not restricted to mongo.
Some people understand the tools they work with. Some people know just barely enough to throw things together and don't tolerate it when something doesn't work out of the box. Worst of all, this second group tends to be very vocal on the interwebs.
I'd almost like to see 10gen not publish the 32-bit package at all. Source is still there. If you want 32-bit, cool, compile it. But forcing the user to compile the 32-bit version assures at least a minimum bound of technical proficiency (an "I understand what I'm doing, why it's not the default and what the limitations are").
I'm a total RDBMS nerd, and it's amazing to me how few people truly care about their data storage. They just want it to work - and, I suppose, it's hard to blame them for that.
*Not that I mean to say that this is the only reason to use a NoSQL DB - it doesn't seem an uncommon one, though.
While you still have to think about your schema, it does mean that you're not constantly writing and removing migrations (rails), while an application is still evolving.
There are trade-offs to each approach but it is probably one of the areas that Rails could still improve by looking at other ORMs - I'd prefer to see the schema specified along with constraints etc for each field at the top of each model to make it explicit and self-documenting, and perhaps doing away with migrations altogether.
So what happens if I have 2 sequential failures? Suppose I have a replica set of size 5 and the master fails? The remaining 4 would elect a new master from amongst themselves, right? But then what if this next master also fails? The remaining 3 nodes are still a quorum (3 > 5/2) and thus (theoretically) should be able to elect a master. But am I to understand that they won't be able to do so?
I'd love to read something describing the "perfect use cases for Mongo" from you :)
Any recommendations for such a tool?
I've also used Munin (http://munin-monitoring.org/, there is a great plugin - https://github.com/erh/mongo-munin), CloudWatch (http://aws.amazon.com/cloudwatch/) and various in-house ones as well.
Other than that, it's descent and free!
http://code.google.com/p/mikoomi/source/browse/#svn%2Fplugin...
Or people are careless about what systems they put into production?
Also, awesome article!
JSON is picky in this regard, and I don't want to convert the whole string to B64 etc encode/decode it going in and out, as I would like to retain regex search capability for the 99% of email titles and names which are not Chinese within mongo from my php application which lives on the front.
http://php.net/manual/en/class.mongobindata.php
You can't do things like regex searches on binary data, but since MongoDB supports different data types within the same "column", you can just store some as UTF8 and some as binary, depending on whether the string has non-UTF8 characters in it.
But the thing I don't understand is, if people use replicasets, how comes they're not using encryption? It would be easy to sniff data off the instances. But yet, when I search on stackoverflow/serverfault, there are close to no people using SSL with Mongo.
I have been using MongoDB for a long time, unfortunately mostly this has been small applications, so you don't really get to test how MongoDB scales.
On that same note, I would love to see a list of gotchas for Riak (assuming some exist). I keep hearing recommendations for Riak, it would be nice to know how it fares in a large production environment.
(1) There's no need to add a "created" field on your documents. You can extract it from the _id field by just taking the first 4 bytes.
(2) If you are storing hashes (md5 for example), you might want to consider storing them as BinData instead of strings. Mongo uses UTF-8 so every character will be at least 8 bits whereas you can get away with 4 bits per character.
Selecting is also reasonably easy, the first 4 bytes of the id are the timestamp (seconds since the epoch). You just create a hex string in that format -- 4 bytes of timestamp and then 8 bytes of zeroes and then create an object ID (using the classes provided by your driver) and do:
coll.find({_id:{$gte:<id>}})
or whatever is the equivalent in your language of choice.And
"Process Limits in Linux" If you experience segfaults under load with MongoDB, you may find it’s beacuse of low or default open files / process limits
Unless I'm out of date, Redis and High Availability don't go together in the same sentence; awesome as it is, it's still a single point of failure.
Clustering is a work in progress (http://redis.io/topics/cluster-spec , http://redis.io/presentation/Redis_Cluster.pdf), replication is available (http://redis.io/topics/replication).