MongoDB 1.7.5 Released with single server durability
groups.google.com
groups.google.com
Do you really want that enabled by default right off the bat?
MongoDBs killer feature was/is sharding. If your deployment isn't going to require sharding from the get go, then I'm not sure why you'd be attracted to it instead of any of the other more mature alternatives.
Adding single server durability just gives MongoDB more possible use cases (e.g. small site, single server, low overhead enviroment, etc.).
I once tried switch to sharding to avoid high write lock ratio(MongoDB use DB level lock currently, only one write operation allowed at the same time). After sharding, MongoDB even don't know how to count my colleciton. db.mycollection.count() return values at random. mongorestore also failed in sharding setup, there's no error message when I was restoring millions of documents, but after it reported successfully restored, I checked the DB, no documents there.
We switched back to more reliable master/slave setup finally. So before trying sharding, you should watch this video http://nosql.mypopescu.com/post/1016320617/mongodb-is-web-sc... . It's so true to some extend.
But if what you say is generally true, then MongoDB is giving a terrible showing for software that's past beta.
I found the following two posts very insightful:
1, http://www.mikealrogers.com/2010/07/mongodb-performance-dura...
"This (mongodb's write plolicy) is kind of like using UDP for data that you care about getting somewhere, it’s theoretically faster but the importance of making sure your data gets somewhere and is accessible is almost always more important."
2, http://www.paperplanes.de/2011/1/10/mongodb_and_data_durabil...
"It's okay to accept trade-offs with whatever database you choose to your own liking. However, in my opinion, the potential of losing all your data when you use kill -9 to stop it should not be one of them, nor should accepting that you always need a slave to achieve any level of durability. The problem is less with the fact that it's MongoDB's current way of doing persistence, it's with people implying that it's a seemingly good choice. I don't accept it as such."
Until 1.7.5, the advice seems to have been that ANY single server is vulnerable, always use replication sets to prevent losing data.
While I appreciate that point, and we do use ReplSets for every DB, in the real world problems happen.
A circuit might explode in a DC, causing all the machines to go down. (Happened to me at ThePlanet)
Our Secondary machine might go down, and while fixing it, the primary might fail. (Happened two weeks ago on dev machines)
Our devs might run a test database on their Macbooks; While this isn't mission critical to stay up, potentially losing records means they need to restart all tests after an event, rather than resuming.
There's a million other places that this will be helpful. Yes, we should always spread things out as much as possible.. But I still use redundant power, RAID arrays, a journaled filesystem and in ideal times a ACID DB.Here's hoping 1.8.0 will be out soon! ;)
And I don't see why 'collection' should take a callback either, unless you're in strict mode and querying mongo to see if it exists.
Some people have written higher level APIs on top of node-mongodb-native, two that I have seen are mongoose and mongous. I'm still evaluating them to see what their performance characteristics look like.
Single server durability is a disk buffered list of pending writes so if the server reboots while in operation it can just resume where it left off. In larger deployments this risk is handled by having multiple servers running concurrently.
Does this sound right?
This covers what MongoDB is doing for durability.
FWIW, here's the FAQ on performance with durability enabled:
How's performance?
Read performance should be the same. Write performance should be very good but there is some overhead over the non-durable version as the journal files must be written. If you find a case where there is a large difference in performance between running with and without --dur, please let us know so we can tune it. Additionally, some performancing tuning enhancements in this area are already queued for v1.8.1 and beyond.
And of course, if you're on SSDs, this is all a non-issue, but that also still comes at a premium.