3 reasons to use MongoDB
ryanangilly.com
ryanangilly.com
Simplified queries, though, are a knock against mongo. Joins are great and I would like to do joins on my Mongo documents, but I end up having to replicate a lot of that in code. Sure it's nice that a document can be more complex and you don't spend a lot of time moving things into tables that are really part of the same record. It's nice because it's not forced, though, not because keeping data in different tables is always the wrong way to do things.
Frankly, I also use MongoDB, and I'm terrified of screwing much with the schema, because then I either a) have to make certain I have anal documentation about what fields are in use on what subsets of objects, keeping code around to make certain to detect and interpret old kinds of data, or b) use "simplified queries" (really, write a bunch of manual code as if I didn't have a query model at all) in order to find and update these old objects, non-atomically and with no transaction safety.
Seriously: the only reason I've so far heard for why having dynamic type verification in my database server is valuable is when you have /so much data/ that it is now fundamentally infeasible to make changes to it with any centralized transaction control--specifically, Google's "it's always the interim somewhere" scenario--not because it is somehow more convenient to do so when you only have in the tens of millions of rows.
1. Don't run javascript on a production db node. db.eval locks the node it's running on until it finishes, so the performance of that node will go down the tubes. Mapreduce is less bad in this regard because it does yield, but it does so too infrequently. If you want to use mongo's built-in javascript interpreter for anything other than development and administration, set up a slave to run your scripts on.
2. Don't use 1.6.1. If you're using 1.6.1 right now, upgrade to 1.6.2. 1.6.1 has a nasty crashing bug that had my mongo node going down about once a day and not coming up without running --repair.
3. Evaluate how much data loss costs you. Mongo stages writes in memory, and so if the db crashes hard it's likely that there will be some data that hasn't made it to disk yet. If you're building a social network the cost of some potential data loss is probably much less than the savings in hardare, admin costs, development costs, etc. But if you're a payment processor or a gambling site, stick with postgres.
This also depends on what is being stored. I would be unhappy if Facebook lost any of my data, and my understanding is that they use safe storage mechanisms (ones where the commit goes all the way to disk before returning) for everything except transient views like the news feed and search. Also, I don't think it's clear that MongoDB has significantly improved either admin or development costs over its safe competition, so we probably only need to look at performance wins.
I'm using mongodb right now and the synergy between jquery - node.js - mongodb is simply amazing.
Users
- id
- events_attended (array of event ids)
Events
- id
- users_attended (array of user ids)
If you want to get events a user attended, do a query like:
1) get the user.events_attended array
2) events.find({_id:{$in: [user.events_attended] }})
3) now you have a nice json object of all events the user has been to
If you want to get all users that went to an event:
1) get the event.users_attended array
2) users.find({_id:{$in: [event.users_attended] }});
3) now you have a nice json object of all users who attended and event
Storing event_ids as an array in the events field defeats the whole purpose of organizing your data into rich documents.
In real life, you very well could build something where the events info was built straight into the user record.
db.events.find({hosted_by: MY_ID})
db.events.find({invited: MY_ID}) // invited is an array of usersIn this case, the events in the user document are events that the user is HOSTING.
There would be another object in our data model to represent the people that got invites to the event. These 'recipients' could be another collection or a list embedded inside each event object.
db.users.find({'events.invites.email': 'sv123@gmail.com'})
and that would return all the users that had invited you to an event.
Now, like I said you're testing my example. When these kinds of requirements are taken into account, you'd probably want to have a separate collection for events. Then you could do:
db.events.find({'invites.email': 'sv123@gmail.com'})
Databases are much better at handling discrete data than file systems - that's what they are built for. Sure, I could keep my data in a bunch of little files, but that doesn't work as well.
(MS SQL has a feature where you "store" the file in the database, but the db writes the file to the filesystem, and just maintains a pointer to the actual file - not a bad hybrid)
I don't know how well GridFS stacks up (it is on my todo list), although I do like the idea of replication and sharding being built in. My gut (which has been wrong before) says that it is good for websites, not so good for general storage.
I use MongoDB for the same reason as mrkurt: prototyping new schemas is a breeze. I still find myself reaching for the old RDBMS toolbox as things move along, grow, and stabilize. Sometimes, a JOIN _is_ the right tool for the job.
What about batch processing a large number of small files? say 10 million image files of 500KB. A typical file system will need to seek each small file.
I wonder if GridFS stores small files in blocks to allow efficient batch retrieval for processing.
The author of the blog post touts it as a _feature_ of MongoDB, but it's more accurate to say that it's an artifact of MongoDB's 4MB document size limit -- you simply cannot store large files in MongoDB without breaking them up. Sure, by splitting files into chunks you can parallelize loading them, but that's about the only advantage.
Among the key-value NoSQL databases, Cassandra and Riak are much better at storing large chunks of data -- neither has a specific limit on the size of objects. I have used both successfully to store assets such as JPEGs, and they are both extremely fast both on reads and on writes.
Neither is built for that purpose, and will load an entire object into memory instead of streaming it, so if you have lots of concurrent queries you will simply run out of memory at some point -- 10 clients each loading a 10MB image at the same time will have the database peak at 100MB at that moment.
Actually, Riak uses dangerously large amounts of memory when just saving a number of large files. I don't know if that's because of Erlang's garbage collector lagging behind, or what; I would be worried about swapping or running out of memory when running it in a production system.
File systems are the most prevalent K-V database.
When the author got to the real arguments, he kept comparing MongoDB to SQL databases and the jab at CouchDB (and the other non-relational databases) seemed without merit to me.
I'm sure there are good reasons to use mongo over couch but I don't think they're the ones listed here.
http://www.mongodb.org/display/DOCS/Tutorial
However I have to respectfully disagree on the "Simple queries" bit. The SQL example given is kind of terrible, however how about-
SELECT * FROM users WHERE id IN (SELECT user_id FROM events WHERE published_at IS NOT NULL)
or
SELECT * FROM users WHERE EXISTS (SELECT 1 FROM events WHERE published_at IS NOT AND user_id = users.id)
(Never use group by as a surrogate for IN/EXISTS. It forces the server to do a lot of unnecessary work)
Is that really unintuitive? Perhaps it's just acclimation, but I find those incredible easy to grep, with the MongoDB example being a variant of the same thing.
Of course that's just for very basic queries. Aggregations in MongoDB are far from intuitive (http://browsertoolkit.com/fault-tolerance.png).
I probably should have thought out the example a little more. We don't ever actually write that kind of query against our production database @Punchbowl. We have a data warehouse pull out high level stats every night, and we query that.
WRT aggregations, you're right -- they do require a bit of acclimation. Once you write a few, though, you're good to go.
Databases like PostgreSQL are excellent at performing joins -- which is all this subselect really is, namely joining two relations -- even when the datasets are quite large.
But this particular MongoDB query comparison is pretty worthless, since it's simply giving an example of denormalization, a concept which is equally applicable to relational databases -- the main difference being that with MongoDB, you hardly have a choice in the matter, since joins don't exist.
Don't get me wrong, I love MongoDB, but there are much better reasons to use MongoDB, such as the fact that every document is a flexible data structure, not a strict collection of columns. You can add keys and values as you choose, and store them as arrays or sub-documents depending on the encapsulation you need, etc.
So generally you will have an easier time working with data and being impulsive about it, than the square-hole-fitting-only-square-pegs model of relational databases, which require more planning and schema design, which in turn tends to squeeze all the fun out of working with databases.
There are pros and cons to both approaches, of course. MongoDB is not as mature as modern relational databases, by far. On the other hand, it has a nice feature which nobody apparently mentions: With MongoDB, the old relational theorist's pet peeve about the meaning of null values becomes moot, because in MongoDB a null value (ie., a missing value) is simply a value which is not there, ie. its key is simply not there. That's much better than null values!
Another advantage is the ability to work with hetereogenous collections of data without having to jump through too many hoops. For example, you can have a collection (table) called "publications". In this table you can store different kinds of publications: Books, magazines, comics, newspapers and so on. Each type of publication may have some common fields, but many have type-specific fields -- hence, hetereogenous data.
A relational database designer will tell you that in the relational world, you would denormalize. A central "publications" table with all the common columns, and then tables "books", "magazines", etc., with each table having their type-specific columns, and also having a foreign-key reference back to the "publications" table. Fine. But think of all the joins you will need just in order to list all and query this stuff; if you have only the publication ID, you have to go through all the tables to determine what type of publication it is. There's not just the performance aspect. The relational model is quite different to how people _think_ about data. MongoDB is easier on the brain, that way.
Yup. I dropped the ball on this one. Should have a list of 4 reasons. :)
A moderately decent database server would have no issues in such a case, however yes, optimization would be case-specific, just as it would be with MongoDB. I was simply comparing readability of SQL for the example that you gave.