MongoDB 2.2.0 Released
mongodb.org
mongodb.org
The new DB level locking introduced in this release is a joke. There's not much difference between that and the old global write lock unless you split your database in dozens of smaller ones. What a pain. I wish they'd just stop pretending and addressed the issue properly once and for all.
I really want to like and use MongoDB because the way data is represented and how it can be queried is awesome.
I benchmarked this at http://blog.serverdensity.com/goodbye-global-lock-mongodb-2-... and there is a video explaining how this works at http://www.10gen.com/presentations/concurrency-internals-mon...
I definitely won't be looking at MongoDB again until this is fixed for good, and until I know replication is more reliable. Had terrible problems that that too unfortunately.
In my experience replication has been extremely stable and reliable for several point releases so I'm not sure what problems you had.
David may use MongoDB at really high scale, but since when is that a reason to call someone an expert.
I don't see anyone nipping at their heels. They have some competition in the key/value space, but there are a few different KV paradigms.
MongoDB is a complete replacement for most RDBMS apps and they are way ahead of the pack.
I'm willing to suck up some inconvenience with MongoDB, it's one of the most inspiring technologies of the last decade. Web application persistance feels more like a fully connected limb now.
I can understand the gripes but do urge people to give 10gen some slack and appreciate that what they are building takes time - it's clear that the team are working hard on features driven by community feedback.
The truth is, MongoDB is awesome if your dataset fits in memory, but if you're in a write-intensive environment with 400-500GB of data, it's just not there yet.
Since you seem to be fairly negative about MongoDB but light on details, perhaps you should write up your experiences so others can learn from what you did wrong.
notes summary:
• Aggregation Framework to fix some map-reduce woes.
• TTL Collections
• DB Level Locking (A step in the right direction)
• Better yielding on page faults
• Tag aware sharding (HELL YES)
• Better Read Prefs
• Indexes now handled by mongodump/mongorestore
• mongooplog replay is awesome for getting point in time backups
• Shell now has full unicode, multiline command history, $EDITOR support (all from change to linenoise.c
* Obligatory unscientific, probably not meaningful, etc. disclaimer.
Mongo Version: 2.2.0-rc1
Hardware: MBP, Snow Leapord, 2.2 GHz Intel Core i7, 8 GB mem
Data: Single collection with 500k records (machine generated time-series event data)
Query Pipeline:
[
{
$match: { ts: { $gte: 1293858000000, $lt: 1296536400000 } }
},
{
$group: {
_id: 'aggregations',
sum: { $sum: '$foo' },
num: { $sum: 1 },
avg: { $avg: '$bar' }
}
}
]
Results: The time range matched against above matches 42,466 documents within the collection. The average response time over 50 runs is 419ms. Not exactly "Big Data OLAP" stuff just yet, but plenty fast enough for most use cases involving reasonably small sets of data. Great job to the MongoDB team!By the way, if you (or anyone else for that matter) come up with useful benchmarks I'd love to get a copy of them at mathias@10gen.com. I have a few of my own, but I'd like to get some real-world workloads from the community to test potential optimizations against.
db.events.ensureIndex({ ts: 1 });
I'll try to clean up my benchmark code a little, throw it in a gist, and then I'll send it your way.The cost, though, is that you wind up having a difficult time doing some things that MongoDB can't do. (For example: Renormalizing your database... Does that even mean anything for MongoDB?)
There's something to be said, of course, for simplifying your design. But it's probably a good idea also to make sure your design reflects your requirements.
You would also have to write a parallel query engine.
I too am a fan of simple designs, but I think rolling your own sharding on top of a RDBMS would likely be a massive chunk of time.
There are really expensive commercial products working on horizontally scaling RDBMS... but personally, I prefer open source and document oriented databases :-)
I like to call that "nice problem to have" territory. There is nothing to fear from being too successful, or from having the problems of success. Success problems can be solved by the application of people, time and money. And like all optimization you measure first to make sure you really have the problems you think you do, Amdahl's law etc.
Far more likely you'll have the problems of not being successful, such as no attention or money from customers and investors. Or not being in the business you think you are in. There isn't really that much point putting in infrastructure just in case you get successful, if that same infrastructure takes time to develop, and slows development.
I'd be delighted having a service "too big" for MongoDB! After all it would mean being more successful than the companies listed at http://www.mongodb.org/display/DOCS/Production+Deployments
There are a lot of proposed work arounds, and some surely work fine if you're only dealing with 2 decimal currency.
The solutions don't scale for arbitrary precision based on the field.
The new aggregation framework is a fantastic step forward and from my testing is relatively peppy, even at a 150k document collection.
Side note: Anyone know a nosql solution that doesn't treat decimals as floats?
I haven't used Mongo in a year. Are Map Reduce jobs still single threaded?
http://www.10gen.com/presentations/concurrency-internals-mon...
http://blog.engineering.kiip.me/post/20988881092/a-year-with...
Almost all of it is still valid. "To be fair, the global write lock is now JUST a DB level write lock. Living in the future guys."
I would like to see ARM support :-)
Of the open source projects you follow, how many still do builds for Solaris? How often do you have to build them yourself?
Not many do. There are quite a few things I build myself if I really really want to use it. A lot of the time it isn't worth it.
I would totally be using MongoDB for some projects (it fits the bill -perfectly-), but my resources for those are usually limited to Solaris on SPARC.
I do have an x86 desktop with Linux as my main PC at work, so I don't miss out on all the fun completely.
However, I do mostly sysadmin stuff, so a lot of the things I use come in the form of scripts, so it's often cross-platform.
Though I do come across some install scripts that just blatantly assume that everything is running Bash on Linux. Those are fun.
I have very little expectation that the MongoDB team will ever push a non x86 version. They do a lot of optimisations deep inside that rely on architecture specific things. But one can hope :)