Does MongoDB 1.3.x (dev) silently lose your data?
korokithakis.net
korokithakis.net
** NOTE: when using MongoDB 32 bit, you are limited to about 2 gigabytes of data
** see http://blog.mongodb.org/post/137788967/32-bit-limitations for more
If you manage to miss that, it's your own fault. Combined with the fact that he was using an unstable development branch, he has absolutely zero room whatsoever to complain. He hit a documented limit and got less-than-friendly behavior on an unstable build of the software. Going nuclear over it is just silly, and makes him look inattentive and/or naive.> makes him look inattentive and/or naive
Oh I wish you gave me more options. This is not a proper ad hominem.
If you want to use Mongo for something serious, make sure you're on a 64-bit box. Memory-mapped file I/O is an important part of what makes Mongo so fast and attractive, but it has obvious limitations on a 32-bit box.
Come to think of it, the article is still a good cautionary tale. What it really teaches is:
a) do your research
b) use the production version if you're using code in production, unless you have a really good reason not to and have some safeguards in place
Also, if the 500k documents limit on 32-bit architectures is true then it's just sad.
Have you ever done software development?
And yes, I did my share of software development...
Is it a "bug" if you redline your car's engine and cause it to throw a rod? Or is that simply you pushing the machine past its design limits, despite the warnings not to do so?
The one thing I found funny is that deleting documents from a database leaves the space allocated and you have to manually run the repair command to reclaim the space, but that requires 100% additional space because it creates a new db and copies the data there.... little bit of a problem since we were already a 99% for the device :-| other than that - been really happy with it.
As of version 8.1 (out since 2005), VACUUMing is an automatic function inside the server. (autovacuum)
The title is incredibly misleading, especially given that this is a ~2 month old post.
As of this writing, 1.4.x is the stable branch, with 1.5.x as the unstable.
I should note that I've been running MongoDB in Production since last August. Development, deployment and go Live occurred on a pre-GA 0.9.x version and I never encountered ANY issues. With almost a year of uptime I've had no data loss or anything else.
At the same time I understood I was using an unstable version and was careful to also understand what was going on under the covers.
So here's the REALLY important thing you should understand if you're using MongoDB because it probably has to do with his data "loss":
Data operations in Mongo, viz. insert/update are asynchronous - from both a API client and a disk-persistence standpoint. One of the things that gives you the speed boost you see from MongoDB is this concept.
When your language driver sends new rows, it does NOT by default wait around to see if Mongo saved it correctly. It sends it off, makes sure MongoDB got the data and returns to you. You can force it to wait, and you can ask it for the lastError - but normally you don't wait for a "I got it in correctly" answer.
Additionally, once MongoDB receives it, it goes into memory... it writes to disk lazily, rather than immediately (if you sent 1000 increment commands between the last disk write and the next, it would batch them into a single increment, etc).
Yes, these things can lead to data "loss" - or the appearance thereof. If you're sending in a bad update or insert statement and not checking for errors, data will "disappear" - by which I mean it was bad data and never accepted for write.
He procrastinated by trying out cool new software (a classic symptom of a project that is not going properly) and lost a lot of time because of the mistaken manner in which he did it.
He should get back on track by focusing on the real meat of his project and use the simplest bitbucket he can find for it.
Oh yeah, hi BJ :)
Besides, whether he was procrastinating or not is irrelevant to his complaint.
[1] ok, some are.
I disagree with your comment. I do machine learning and NLP, and I prefer mongodb as a data-store to RDBMs. SQLite would be overkill. MongoDB is a more convenient and simple choice than an RDBM.
Using a schema-less document database is natural in NLP, where you might add columns to your fields on the fly, e.g. you are experimenting with a new preprocessing step and want to cache the result.