MongoDB 1.8 (stable) released
blog.mongodb.org
blog.mongodb.org
I want my data store to be durable and unsurprising -- barring a hardware failure or such, if I submit data it should either tell me that it failed to commit or it should be stored durably and without surprises (e.g., it should not truncate a long string to fit).
I've read some of the Mongo docco, and it's pretty exciting, but the lack of ACID -- primarily the Durability -- has kept me from really using it.
With a WAL journal, it sounds like maybe the durability issue is fixed. Is it? Could I use Mongo with relatively out-of-the-box settings plus --journal and count on a level of durability equivalent to a traditional RDBMS?
For example in the PHP driver, calling insert with the safe option. http://www.php.net/manual/en/mongocollection.insert.php
"If safe is an integer, will replicate the insert to that many machines before returning success (or throw an exception if the replication times out, see wtimeout)."
You can then immediately called http://www.php.net/manual/en/mongodb.lasterror.php to confirm the last operation didn't error.
The mailing list is a great place for this type of question, too (http://groups.google.com/group/mongodb-user).
"New map/reduce options for incremental updates" would also be really cool if they had a way to do something like couchDBs incremental views. This would require keeping track of changes or a "trigger" functionality that runs the m/r task after every x inserts
EDIT: Also... does the group commit mean that ALL write transactions will be un-acknowledged to the client until the group commit finishes?
They'll really need to add some major smarts into the journaling and group commit if they're going to be able to stack up concurrent I/Os to help feed big disk arrays and even to get the best use out of SSDs which make back-and-forth latency on the I/O pipes even more significant.
- horizontal scalability
- flexible datastructures
- map/reduce
If you're interested in polygons, lines, etc; more physically accurate (and completely implemented) distance queries, spatial joins, aggregation, 3D and surveyor-annotated data, set-theoretic operations ... PostGIS is far and away the way to go. It's far more mature and debugged than any of the NoSQL geospatial stuff I've seen, not only WRT correctness but also performance.
As a point of reference: There's a growing legion of geographers who do all their vector work in SQL using PostGIS.
All that said, for some applications, being tied to the relational model is a deal breaker. Just know that in terms of capability and maturity on the geospatial front, you'll be trading off a Cadillac for a partially assembled rocket sled.
> We don't currently handle wrapping at the poles or at the transition from -180° to +180° longitude, however we detect when a search would wrap and raise an error.
generalized grumble
Why does everyone always seem to punt on doing geospatial right? It's not _that_ hard.
(It's a nice opportunity to publicly show off a specialty/core competency and brush up a bit on C++ a the same time. I'm not that easily provoked into action by internet commentary! ;) )
But, looking at the source, I think I will be probably a Bad Contributor and end up with a gigantic pull request and a (mostly) full re-implementation...
Do you mean you think they don't know how to do it?
As I followed the roadmap on this specific point, it looks more like an incremental development to me: they first used rectangular coordinates in 1.6, then a spherical model in 1.7 etc.
It allows to bring a more lightweight solution quickly to people that need it (like me), then to evolve based on the feedback etc.
For example: If they were truly using a "spherical model", then one would not expect to have queries fail at the poles & dateline, would you?
At least it is documented and fails hard with an error rather than giving wrong results, so a developer can quickly figure out the weak spots --- though I bet a lot of people would prefer the wrong results to queries that cause exceptions in their systems.
> Do you mean you think they don't know how to do it?
I think it has more to do with the absurdly low bar they've set for themselves to check the "geospatial" box than it does with competence.
Your mileage may vary as they say: I used the GIS since 1.6 and it was very helpful for me in this form already :)
Analogy: It's like seeing "ACID compliance!" on a feature list, then finding buried in the documentation that is only the case for single-document transactions in unordered collections on a single machine only.
The new feature might be useful to some but including it on a feature list without disclaimer is misleading.
Really curious: are you using mongo currently? Or browsing the docs?
It's very fast and the flexible schema makes the code much more flexible and easy to write. And did i mention it was fast?
You should definitely give it a try and consider using it for such systems.
My only issue with it is that i am running it on a 32-bit system and so i'm limited to 2GB a database.
Initial attempts in SQL were painful. The only real way to do it was a key value table, but that gets painful when it comes to formatting for web presence (notably, each document has sections with a group of fields, plus some fields may need to be grouped together such as a series of checkboxes, or parts of a name). So at that point we're looking at writing up XML files to describe the presentation of these 200 forms from a key/value table to the web app.
At that point I realized this was doable, but going to be a mess. Enter mongo. Mongo essentially let's us store a dynamic schema of documents. For each form we can stick it all in a single document, as a series of embedded models, with all metadata and values needed in one go. We also get nice revision control within that using mongoid. We can now fetch all the data for a form, as well as save all the data for the form, in one VERY fast atomic operation (we're talking 100-800 field definitions for each form). Having never used mongo, it only took me a few days to implement this complete with handling for all field types and performance was fantastic.
Mongo also made it quite easy to populate our data since we're essentially just storing a tree of key and values. We wrote up a tool that loads up the PDF's and let's us draw boxes on top of the fields and set up the metadata, then export that to a YAML file for each form. The YAML is then stored in a tradiational SQL database and is used to create a new form in the system by simply converting it to a nested hash and having mongoid save it. Slick.
I'm getting a bit wordy here, but I think it's a great real world example of the type of problem mongo is a good fit for. I wouldn't personally use mongo for something that a relational database is a good fit for, but for something like this it allows you to solve the problem quicker and with significantly less code to maintain (really, the CRUD code for forms is no more than with SQL and probably less since it's only one operation on a document, and my pdf form generator is < 200 lines of ruby).
This saves a lot of time you'd normally spent defining schemas, and is very flexible. It didn't completely replace SQL for me, but it's a good fit for the heterogeneous free-form data generally encountered on the web.
I work a lot on data aggregation, where I can create a bunch of tables each day, then maintain them etc. For me it's almost a dream really :)
As well they are adding features such as geonear that makes it appealing for other uses (which I have, too).
This is Not exciting:
"Note that this option is possible only when the result set fits within the 16MB limit of a single document."
But you're right too, of course.
http://blog.evilmonkeylabs.com/2011/01/27/MongoDB-1_8-MapRed...
(Disclaimer, I work for 10gen / MongoDB)
You guys do seem to be headed in the right direction technically... I just can't bring myself to say "mongo" out loud.
http://groups.google.com/group/mongodb-user/browse_thread/th...
See http://news.ycombinator.com/item?id=2052852 for a comparison.
Redis is a great key-value store, MongoDB is more of a fully-featured database. Redis has some nice set operations and is pretty easy to learn (all of the commands are here: http://redis.io/commands). MongoDB is also pretty easy to learn (click the "Try it out" button at http://mongodb.org/), but there are a lot of advanced features to learn about.
So, if you need a key-value store, Redis is a great choice. If you want to do something more complex, MongoDB would probably work better.
In both cases, (unlike CouchDB) you can alter data structures by more complex means than simply replacing the whole thing (such as incrementing a counter). In both (again unlike CouchDB) the updates overwrite in place and do not waste space (but also do not preserve past versions or allow readers to overlap writers).
Redis is for stuff that fits in memory. MongoDB scales up to "big data", provided the individual items are moderately sized.
Redis runs in RAM so it's blazingly fast. MongoDB is about as fast as MySQL.
Redis is single threaded so only one operation runs at once (the speed makes this mostly not a problem). Some operations globally block MongoDB, some can run in parallel.
In both, operations are atomic. Redis has transactions of a sort that group operations and ensure the data they relate to is unchanged. MongoDB operations can't be grouped into a transaction, but they can be a lot more complex so they effectively become a transaction (limited to operating on one data item).