You buy a car. It comes with brakes disabled because for whatever reasons that also lets it get to a higher top speed. You are expected to read you car owner manual and on page 54 you find that you have to hold "enable brakes" button under the console for 10 seconds to turn on your brakes. Would it vex you that people might be slightly critical of that car. Clearly they are silly for not reading their car manual until page 54.
That "feature" is not something that should be discovered by reading docs or when you get a crash and then load a backup from another week and still get a crash and then you start hitting your head on your desk.
Anything calling itself a "database" should not have shipped with those default settings _ever_. If they did they might have gotten away with it in my book by having a big flashing red warning on the front or download page. I don't remember one.
Don't get me wrong, I'm not saying your assumption is unreasonable. But in the end, it's on you as a conscientious developer to read the documentation. I'm not even suggesting cover to cover - in this case though they are very up front about write concerns. There is no real excuse to find this out any other way, it's just negligence.
It's a nice hypothetical, and that might be your style. Most real-world car purchase scenarios I'm familiar with would make that style impractical.
So, let me get this straight: You laid down tens of thousands of dollars on a vehicle that you only post-purchase read the manual of, and you're raising this as some sort of standard people should follow?
Honestly, asking the right questions (and test-driving) upfront should be what lands the purchase, and not discovering the folly of purchasing a car with such ass-backwards issues you only discover after the fact when you bother to dig out the manual.
You drove it off the lot after you bought it, right? Or did you read the manual in the lot right after signing the papers locking you into the purchase?
Honestly, how can people on HN actually be this against reading? Especially things that are really important? Sure, don't read the contest rules for your McDonald's monopoly. But if the data for your livelihood depends on something, there's no excuse for not reading the documentation.
It's impractical because people live a finite amount of time and this is a terrible use of it
You don't need to read every word of everything, but some things are worth it. Do you sign contracts without reading them too since it's a "terrible use of" your time?
Yes - I am saying in this particular car example, the benefit derived reading the entire manual before purchasing/driving a car is not worth the cost unless your time is worth very little. As others have pointed out, no one is flipping through the manual to check whether the brake pedal actually applies the brakes
But it is a good engineering decision to thoroughly read the docs before jumping into a new datastore like Mongo, I agree. Learning there are things like gigantic global locks and unsafe writes are normally enough to make you say, "hey, I probably shouldn't use this to store production data I actually care about"
A database that has default configured that ends up corrupting users' data silently is like buying a car with the brakes disabled.
Well except that in the car case may brakes disabled won't make the car go faster, but in case of MongoDB I remember fans strutting write benchmarks around comparing it to Postgres, Couch and other database and telling how it is webscale. The reason that design decision was made is shady. That was my initial point.
I did experience one once though: I discovered that a vehicle had traction control when the system activated during a skid. The computer and I disagreed about the best way to respond, and the surprise did make the situation more dangerous than it could have been.
In both vehicles and databases, the situations in which the product might do something unexpected and dangerous should be clearly documented in their own section of the manual. Databases should say "here are the things that could lead to data corruption or loss". Vehicles should say "here are the situations where the vehicle might disregard or override the driver's control inputs".
I suspect some people will believe the results wouldn't have been what I expected without the traction control. I can't prove they would have been, but I did grow up and learn to drive in Alaska. Based on my experience, I think I would have done better than the computer did.
Here's somebody describing the basic idea involved while discussing the joys of mildly irresponsible driving on cloverleafs: http://www.scottgood.com/jsg/blog.nsf/d6plinks/SGOD-66RJ2Y
Heck, even when I have a rental car for a single day, I will read the manual. Maybe not every page, but I will skim it for gotchas (and if I have time, the whole thing).
Maybe it's just the engineer in me, but it's what I do.
Come on guys, as computer engineers/programmers/developers/whatever, surely professional pride at least would mean we at least read the README and/or the manual, before putting something into production?
it's more like launching a shuttle mission
Given how much we spend time talking about MVPs, Lean Startup, etc, I think people on this site are trying to avoid launch a shuttle mission. [Often] They're looking at building startups and are looking for both time-tested and new-but-advantage-providing technologies and techniques. At first glance, MongoDB appears to be advantage-providing so people adopted it quickly. They didn't read the manual; they put it in production on a small site and got surprised by the lack of durability.I use MongoDB now in production and I am happy about it. Not huge dataset by any means so MongoDB fits the bill perfectly. It has a few idiosyncrasies (doesn't release disk space after deleting records - what?!?) and you definitely want to read the manual on settings. But it is incredibly easy to use (documents instead of relational data) and allows me to focus on app development instead of my storage backend.
HBase needs hflush to make sure that the WAL edits are resident at at least 3 (default) HDFS data node machines.
Not sure how exactly grand parent lost data. Each edit is first written to the WAL then committed to the in memory store. The in memory store is flushed to disk into a new file at a certain size. If a server crashes and had unflushed data in the memory store that part of the data is replayed from the WAL on another server.
See also here: http://hadoop-hbase.blogspot.com/2012/05/hbase-hdfs-and-dura...
These are systems designed for the real world, where people don't read the manual until they have to.
When people assume MongoDB was similarly designed with their best interests in mind, that's when things go wrong.
Any time I deploy something as critical as a database, I carefully read about what it does and how it works. Not doing so is like signing a contract without reading it.
I don't understand this reasoning. We are talking about defaults. Defaults are used by people who did not tweak the settings yet. If I am just starting building a thing, I will have bugs and squeaks and I want to make sure I am not fooled by some unreliable data store. I am not likely to need 100GiB/s throughput, but I am very likely to have to hunt bugs, like "I did click on this <like> button but it did not add to the total likes". And I would really really hate it if after half a day of bug hunting I would realize that my data store just didn't store the thing...
No, I just assume that a database has a similar set of features as other databases have had for decades. Mongo does not; it is clearly the exception - and for possibly nefarious reasons, as well.
I don't have much sympathy for people who can't RTFM but storing data is kind of a thing for databases.
if you're not trying for that standard at all, it's false advertising.