From a developer:
“As long as the OS & the HW doesn't crash, the data is safe thanks to the page cache”
This is so strange to me. A database that is non-durable by default. OK…
From a developer:
“As long as the OS & the HW doesn't crash, the data is safe thanks to the page cache”
This is so strange to me. A database that is non-durable by default. OK…
I still don't have an answer for him. It sounds just as strange to me too.
This reflection [1] came from the founders of RethinkDB, a competitor of MongoDB at the time:
"It turned out that correctness, simplicity of the interface, and consistency are the wrong metrics of goodness for most users. The majority of users wanted these three trade-offs instead:
- A use case. We set out to build a good database system, but users wanted a good way to do X (e.g. a good way to store JSON documents from hapi, a good way to store and analyze logs, a good way to create reports, etc.).
- Timely arrival. They wanted the product to actually exist when they needed it, not three years later.
- Palpable speed [...]. MongoDB mastered these workloads brilliantly, while we fought the losing battle of educating the market."
MongoDB narrowed things down for a specific use case, and became the best for that use case. This comes with trade-offs. MongoDB was probably not the best database for healthcare back in the days, but that is OK. It did the job very well for other use cases and industries. And over time, they fixed the issue around losing data and became more stable. Essentially, they made developers feel like superheroes, and over time improved their product, and eventually grabbed a massive market share.
[1] https://www.defmacro.org/2017/01/18/why-rethinkdb-failed.htm...
It used to be open source. It's not anymore
> The parent company is a listed company and worth $15BN, 3x more than Elastic to put some perspective.
That's purely a capitalistic argument and makes no difference to whether the product is any good. For example, there's plenty of "churches" that are richer than MongoDB Inc. and absolutely abhorrent and evil.
> This comes with trade-offs.
The only thing that required the trade-off of data loss was cheating in benchmarks in order to hoodwink naive potential users into using their dangerous product. MongoDB Inc. has always preferred to lie to their users. It is not a database company; it's a marketing company with a product they label as a database. And that's a smart way to make money, sure, because of vendor lock-in, but it's not a smart way to gain trust.
As engineers we bear responsibility for how our work impacts society. Mongodb may have made their investors a lot of money, but they did sloppy work and didn’t do right by their customers. That’s not a success in my book.
There are many cases, and ingest is one of them, where no being durable is fine. If you can either:
* Repeat the whole process on failure (which is assumed to be rare) * Recover from the failure without data corruption (distinct from data loss, mind)
In those cases, being 10x faster is very compelling.
Note that this is about ingest for bulk loads, while online transactions not being durable is a really bad idea.
For bulk load ingest, you can usually retry the whole operation. Not so for transactions.
For some reason, many databases that overwrite data support disabling journaling and/or fsync. E.g. SQLite has "pragma journal = off". You can lose the entire database from an ill-timed crash, if one important page gets written but another doesn't. To their credit, it's not the default, and the documentation is explicit about this:
> If the application crashes in the middle of a transaction when the OFF journaling mode is set, then the database file will very likely go corrupt.