Jepsen: MongoDB 4.2.6
jepsen.io
jepsen.io
In mongo, you have a `withTransaction(fn)` helper that passes a session parameter. Mongo can call this function mutliple times with the same session object.
This means that if you have an async function with reference to a session and a transaction gets retried - you very often get "part of one attempt + some parts of another" committed.
We had to write a ton of logic around their poor implementation and I was shocked to see the code underneath.
It was just such a stark contrast to products that I worked with before that generally "just worked" like postgres, elasticsearch or redis. Even tools people joke about a lot like mysql never gave me this sort of data corruption.
Edit: I was kind of angry when writing this so I didn't provide a source and I'm a bit surprised this go so many upvotes without a source (I guess this community is more trusting than I assumed :] ). Anyway for good measure and to behave the way I'd like others to when making such accusations here is where they pass the same session object to the transacton https://github.com/mongodb/node-mongodb-native/blob/e5b762c6... (follow from withTransaction in that file) - I can add examples of code easily introducing the above mentioned bug if people are interested.
I just don't want to be called to the office on a weekend anymore for this sort of BS.
Production incidents with MongoDB last year: 15 Production instances with redis, elasticsearch and mysql combined last year: 2 (and with much less severity)
Edit: just to add: I didn't pick Mongo, I was just the engineer called to clean that mess. I created enough of my own messes to not resent the person who made that shot for it. We are constantly on the verge of rewriting the MongoDB stuff since a database that small (~250GB) should really not have these many issues (In previous workplaces I ran ~10TB PostgreSQL deployments with much more complicated schemas and queries with far fewer issues). It's also expensive and support at Mongo Atlas hasn't been great (we should probably self host but I am not used to small databases being so problematic)
aw crap. oh well it probably doesn't matter for my small-ish application.
I suppose salespeople probably aren't into the nitty-gritty, but their tech people should have warned them about this. Maybe they were just trying to pull our collective leg, but I suppose that why I was at that meeting.
It was obviously an instant 'No'.
Even if we weren't - as a sales engineer on a large CMS/ECommerce platform with merchants running $150M+ in annual revenue, with an average client retention of seven years, and two decades of agency experience behind the decisions around building that platform, if you instantly said no just because of MongoDB, maybe you don't know as much about MongoDB as you think you do.
I came from a SQL background myself, and had reservations based on all the things I'd read about MongoDB as we decided to build a platform after doing things bespoke for two decades, but time has proven our architecture choices out. It's easy to be proud of something that works well.
My only experience with MongoDB is being "the engineer called to clean the mess". I'm sure you can effectively use MongoDB in production if you're knowledgable and careful, but most people aren't and they shouldn't have to know the detailed inner working to not create a mess.
It’s always the same
1. Newbie webdev (aren’t they all) uses MongoDB because it’s easy to use according to blogs and twitter
2. Somehow it makes it into production
3. A dozen experienced engineers spend years trying to keep it running
> Please respond to the strongest plausible interpretation of what someone says, not a weaker one that's easier to criticize. Assume good faith.
[1]: https://news.ycombinator.com/newsguidelines.html
In this case, the parent commenter probably meant that "newbie web developers" are likely to choose MongoDB. Of course, web developers have a range of experience, some new, some seasoned.
Regarding your comment, I am reminded of a pattern in online behavior over time: people seem to take offense more easily. Please take a look at https://www.psychologytoday.com/us/blog/how-do-life/201410/t... to understand what I mean.
Based on seeing how comments like this may get interpreted, as well as broader thinking about online communication, I think HN should consider a more nuanced system of comment feedback mechanisms.
I don't have a particular plan finalized, but I would like to see HN provide feedback on different aspects of the comment. Below are some important aspects:
To what degree does the reader / voter... *
1. agree/disagee with the comment?
2. find the comment relevant / irrelevant to the topic as a whole?
3. find the comment is situated in the correct / incorrect location in the thread? (e.g. responding to the parent comment or not) 4. find the comment interesting / uninteresting?
5. think the comment adds to a diversity of perspectives?
6. find the comment clear / unclear?
7. think the comment aligns / (does not align) with the HN Guidelines: https://news.ycombinator.com/newsguidelines.html
8. find the comment welcoming / offensive?
* When I write '/' above, I intend it to be a continuum; e.g. hot/cold means "in the continuum between hot and cold).
Additionally, being able to give feedback in a more granular fashion could be of use. For example, in my comment above, I would not be surprised if a significant number of people were bothered/offended by my commentary that people seem to be taking offense more easily. Some would call this ironic -- I wouldn't -- I think it gives more data to prove the point.
Motivations: my goal here is not to gain or lose karma -- I care very little about karma here, precisely because it is so muddled and varied from person to person -- as long as I have enough to participate fully. My goal is to learn and play a small part in fostering awareness and community, while hopefully to motivating others to reflect on their impact on the community here.
> "Newbie webdev (aren’t they all) uses MongoDB because it’s easy to use according to blogs and twitter"
I'm not sure if I completely missed the "aren't they all" part or if the comment got edited.*
So, clearly, in response to "aren't they all", the response by sophiebits above makes complete sense.
* This makes me wonder, is a comment locked for editing once a reply gets added?
The thing is that MySQL is older so it went through all of it earlier, but it still suffers from poor decisions from the past. This is contrasting with PostgreSQL, where correctness and reliability was #1 from the beginning. It started as an awfully slow database, but performance for improved and we now have correct, reliable and fast database.
Transaction ID wraparound: https://twitter.com/bcantrill/status/1110647418008133632
Incorrect use of fsync: https://news.ycombinator.com/item?id=19119991
If you were around back in the day you will remember the MySQL team claiming that no one needed transactions or referential integrity, that you should just do it yourself in the application...
ORMs and ActiveRecord in particular encourage, to some extent, the use of a RDBMS, even if they didn't get to take advantage of them well for a long time - for example, in RoR "has_one / has_many" for foreign-key relationship, .joins(:field_name) for, well, joins, and so on.
A big reason I called out RoR is that back in '04-05 I was railing against its default use of plural table names, and DHH on IRC recommended I shut up and just flip the configuration switch and turn off the feature, but of course when I did that all sorts of latent bugs were exposed.
RoR was the beginning of hipster "coding" and I therefore blame it for everything.
I'm wasn't previously familiar with the FriendFeed approach to database (ab)use. I paid about as much attention to it as I did to MySpace back in the day -- nearly zilch -- so its etc innards are doubly obscure to me.
> MySQL was the perfect database for people who had no respect for databases.
Very few database systems support online DDL, which unlike a transaction, does not require undo or rollback resources. Of course one must have a rollback procedure if something fails, but you need one for transactions too, just in case.
An online rollback is far lest costly than a transactional rollback, because and online rollback is just undoing what you did. Added a column you didn't want in one query? Remove it again in another, very quickly.
TokuDB (a mysql/mariadb storage engine) supported all DDL as an online operation. But percona killed it in favour of TokuMX, the MongoDB equivalent.
TokuMX has no upgrade path to wired tiger, only one major customer at Percona (I can't say who it is) and no engineers.
Any kind of DDL is tricky and requires users to RTFM for the intricacies of their chosen database. One size rarely fits all.
I really wish Percona would reconsider their decision to deprecate it.
After Percona took over TokuDB's creator TokuTek, they wasted so much of their development time and money on TokuMX (Percona's fractal tree-enabled MongoDB server) only to abandon it in 2017.
That money would have been better spent on TokuDB development to allow it to match the features present in InnoDB like generated columns, spatial indexing, fulltext Indexes and Galera.
TokuDB still has many users and MyRocks is just no substitute.
People rightfully joked about MySQL when they had the non-ACID engine.
Same for MongoDB. A database that loses data when properly used is a joke.
Yes, there are use cases out there for fast non-guaranteed writes. No, 99% of companies don’t have them.
Periodic snapshots of state held elsewhere. If you lose a write, you just get stale data until the next update.
Firm realtime work. If you lose a write, that sucks, but a slow write sucks just as much.
I could think of an instance where you’d like to log data, but the occasional datapoint being missing wouldn’t be terrible. Maybe something like a temperature monitor — you’d like to have a record of the temperature by the minute, but if a few records dropped out, you’d be able to guess the missing values from context. Something like the data monitoring equivalent of UDP vs TCP.
You then "checkpoint" when the game is over.
You might dissent that is not a "non-guaranted" write, because in fact the write did occur, but I simply want to allude to the concept of a "non-secured" write, in that it vanished without an fsync.
I work for a telco where we log large amounts of network requests using MongoDB.
Also would depend on how long time you think the application will be around. You're building an MVP to evaluate something? Just hack together whatever will work (then throw away). You're maintain software for a library/archive that will most likely stick around for a long time, even if they say it's just temporary? Do decisions that will help in the future, always.
We have a roadmap we need to meet and so far we have been trying to spill money on it rather than developers (paying mongo atlas) and adding features incrementally as Mongo gets them (like transactions).
If this wasn't a startup we would probably rewrite.
That's about a decade out of date at this point. MySQL/InnoDB is the standard table engine and corruption is exceedingly rare. As of 2014, when I last directly worked on MySQL prod systems, there was no practical difference from PostgreSQL in terms of transactional guarantees. That includes APIs like JDBC which we used for billions of transactions.
MariaDB [test]> create table test ( i int );
Query OK, 0 rows affected (0.06 sec)
MariaDB [test]> insert into test values (''), ('xxx');
Query OK, 2 row affected, 2 warning (0.01 sec)
MariaDB [test]> select * from test;
+------+
| i |
+------+
| 0 |
| 0 |
+------+
2 row in set (0.01 sec)
There's a bunch of other similar caveats as well, and this can really take you by surprise. I've seen it introduce data integrity issues more than once.That's a new MariaDB 15.1 with the default settings I just installed the other day to test some WordPress stuff. I know there are warnings, and that you can configure this by adding STRICT_ALL_TABLES to SQL_MODE, but IMO it's a dangerous default.
This is also an issue with using MongoDB as a generic database: every time I've seen it used there were these kind of data integrity issues: sometimes minor, sometimes brining everything down. Jepsen reports aside, this alone should make people double-check if they really want or need MongoDB, because turns out that most of the time you don't really want this.
https://wiki.postgresql.org/wiki/Transactional_DDL_in_Postgr...
That MySQL autocommits is also even worse than just "doesn't support it."
A basic element of the relational model is that metadata is stored as relational data and that the same guarantees that apply to manipulating main data in the database apply to manipulating the schema metadata.
It's true that many real relational databases compromise on this element in various ways at times, but it is absolutely not the case that DDL “is supposed to be” non-transactional.
isn't that the point? you can use a session to do multi actions within that session.
withTransaction(async session => {
await Promise.all([someOp(sesson),
someOtherOp(session)]);
});
Mongo may retry running it (calling the function again) if a "TransaientTransactionError" is raised (the transaction is retried from the client side rather than at the cluster).However, when the driver calls your function again it doesn't invalidate the `session` object - so previous calls to the same function can make updates to the database.
Let's say `someOp` does something that causes the transaction to retry and `someOtherOp` is doing something non-mongo-related in the meantime (like pulling a value from redis). Now `someOtherOp` reached the mongo part of its code and it is executing it happily with the same session object (so operations succeed although they really shouldn't)
The point of transactions like you said is to perform multiple operations atomically and for them to happen "exactly once or not at all". With Mongo in practice it is very easy to get "Once and some leftovers from a previous attempt".
What exactly would invalidating that session object do here? And what would the session object do after it was invalidated?
Because you wrap the DB operations inside a Promise.all, it means it will run them all BUT it will not revert them if one fails (it's not atomic, it just says that one has failed and you need to catch it), it will reject them but not revert them. (the CUD operation will already have changed the data) The problem I believe is the transaction is considering the Promise.all and not what's inside of it so it will run it again despite the fact that some have already succeeded earlier
I think you just have to resolve each of them outside a Promise.all. In your case because Promise.all has been rejected it will redo the transaction, therefor it will redo the one that have already worked in the first call.
I'm no expert but this is how I understand it.
Without promise.all, I think it can be replicated like this:
try{
await someOp(session);
catch { }
await someOtherOp(session);
He expected what the session to be invalidated during someOtherOp.They implemented a "disposer" incorrectly (I added this to several other libraries and made a Q&A about it here https://stackoverflow.com/questions/28915677/what-is-the-pro... )
Just wondering, did you submit a bug report to them about this? If so, any response?
We reported it immediately at the highest severity and we pay for the highest tier or support - we tried to collaborate as soon as possible. It sort of went "over their head".
For example at #00636235 "how to avoid TransientTransactionError".
That is, a table on the jepsen.io frontpage, or at least on each product's review page, with database products and configuration on rows and consistency properties on columns, and a nice "Yay!" or "Nope!" mark in the cell, plus links on how to achieve the database configurations in the table (esp. how to configure each database to have the most guarantees).
Also, ideally the analyses should be rerun automatically (or possibly after being paid, but making it easy for the company to do so) every time a new major release happens rather than being done once and then being stale.
Finally, there should be tests for the non-broken databases (PostgreSQL for instance, both in single-server mode, deployed with Stolon on Kubernetes and using the multimaster projects) as well to confirm they actually work.
This is a wonderful idea, and I've got no idea how to actually do it in a standardized, rigorous way. Vendor claims are often contradictory, it's hard to get a good idea of anomaly frequency, availability is... a rabbithole, and it's hard to come up with a standard taxonomy of anomalies--most of the analyses I do wind up finding something I've never really seen before, haha. With that in mind, I've wound up letting the reports speak for themselves.
Also, ideally the analyses should be rerun automatically (or possibly after being paid, but making it easy for the company to do so) every time a new major release happens rather than being done once and then being stale.
I don't know a good way to do this either. Each report is typically the product of months of experimental work; it's not like Jepsen is a pass-fail test suite that gives immediately accurate results. There is, unfortunately, a lot of subtle interpretive work that goes into figuring out if a test is doing something meaningful, and a lot of that work needs to be repeated on each test run. Think, like... staring at the logs and noticing that a certain class of exception is being caught more often than you might have expected, and realizing that a certain type of transaction now triggers a new conflict detection mechanism which causes higher probabilities of aborts; those aborts reduce the frequency with which you can observe database state, allowing a race condition to go un-noticed. That kinda thing.
If I'm lucky and the API/setup process haven't changed, I can re-run an analysis in about a week or so. If I'm unlucky, there's been drift in the OS, setup process, APIs, client libraries, error handling, etc. It's not uncommon for a repeat analysis to take months. :-(
It would be kinda like you including this sort of thing on your resume. Which would also be a bad idea.
https://web.hypothes.is/about/ or similar could be used to develop commentary overlays on top of marketing materials.
Plus maybe a column indicating what [the company behind the database] claims?
Postgres is widely understood to be a robust database with safe defaults. I, and perhaps others, would love to see you aim your array of weapons at Postgres. Do you have any plans to look at stock Postgres?
> Postgres has both asynchronous (the default) and synchronous replication options, neither of which offers automatic failure detection and failover [12]. The synchronous replication only waits for durability on one additional node, regardless of how many nodes exist [13]. Additionally, Postgres allows one to tune these durability behaviors at the user level. When reading from a node, there is no way to specify the durability or recency of the data read. A query may return data that is subsequently lost. Additionally, Postgres does not guarantee clients can read their own writes across nodes.
> > Postgres has both asynchronous (the default) and synchronous replication options, neither of which offers automatic failure detection and failover [12]. The synchronous replication only waits for durability on one additional node, regardless of how many nodes exist [13]. Additionally, Postgres allows one to tune these durability behaviors at the user level. When reading from a node, there is no way to specify the durability or recency of the data read. A query may return data that is subsequently lost. Additionally, Postgres does not guarantee clients can read their own writes across nodes.
> From http://www.vldb.org/pvldb/vol12/p2071-schultz.pdf
This is like those commonly seen tables comparing your product with others where your product had checkmarks in all categories, and of course competitors are missing a bunch of them. The problem is that the categories were picked by you, and are often irrelevant to the other product. This is the case here.
PostgreSQL is not a distributed database, the master is the one doing all writes. The replicas are read only. By default replicas are asynchronous which means they won't affect master performance, at the cost of having data there being late by few seconds. Since you can't write to replicas, this won't cause data corruption, only delay which often is acceptable. If you design your applications in such way that will have two database endpoints: one for writes and one just for reads, you can then decide based on context which endpoint you want to use. The read only is easy to scale, but as mentioned earlier it is read only, and might slight delay.
Now, for failover, you might also opt on using synchronous replicas this will add extra latency, but then you always have at least one machine that has the same data. They mentioned that if you have multiple synchronous standbys then it only one needs to write. Actually that's configurable, you can specify group of synchronous machines and how many and which need to be synchronized, the remaining ones are a backup in case those that you specified aren't available.
Besides, the writes don't work the same way as in mongo, when a standby node is in sync it isn't just in sync for that particular write, it is completely in sync, so their following argument about not being able to specify durability/recency of data on read is redundant. If you contact the master or synchronous replica, you will always get the most recent state. If you don't mind slight delay you should query asynchronous replicas (in fact you should prefer them whenever you can, since those are cheap to add)
> the master is the one doing all writes. The replicas are read only. By default replicas are asynchronous
The same is true with MongoDB's defaults in an unsharded cluster.
I believe RDS Postgres is probably the right answer for lots of applications, especially for those that already depend on AWS for baseline availability. I'd love to see if that holds up against a rigorous analysis.
With tools like repmgr it is just a single command invoked on the standby.
If you absolutely don't want to lose any data, you should have two masters in close proximity (so the latency isn't high) set up with synchronous replication, then have one or two standbys with asynchronous replication. This will reduce throughout, but then you can be sure that the other machine has all the same transactions. If something happens to both you then can fallback to the asynchronous one which might be a bit behind.
Automatic failover for PostgreSQL works great and can be done safely if combined with synchronous replication.
Multiple tools will implement this correctly:
https://patroni.readthedocs.io/en/latest/replication_modes.h... https://github.com/sorintlab/stolon/blob/master/doc/syncrepl...
Quoting a former colleague here, but "if it hurts, do it more often". That is what you should do with your PostgreSQL failovers.
I have clusters running on timelines in the hundreds without a byte of data loss due to using synchronous replication, tools that help out with leader election, and just doing it often.
I would actually be interested if aphyr's analysis of Patroni and other distributed add-ons to PostgreSQL.
The only question is how soon are you going to page humans. After the automated mechanism flipped your master 2-3 times but the cluster still hasn't made progress [nothing coming out of the master; or it locks up after a few minutes again]), or right after some other automated mechanism detects that there's a problem.
Whatever automation you have in place, it has advantages and disadvantages. In the GitHub case - I suppose - they determined post-mortem that it would have been better to just let the master chug through the incoming onslaught of queries instead of failing over, and over, and over. (But of course this seems like a trivial problem in any auto failover setup, so I suspect there's more to the story.)
No. But the contract Patroni has is this:
I only serve a master (primary) if I have the lock. If I do not have the lock I will demote.
This results in that there can be only 1 primary active at any given point in time, even if the network is partitioned.
This in and of itself does not guarantee no-split-brain situations, a split-brain can occur if writes were made on the former primary, but not yet on the future primary. This however can be mitigated with synchronous replication.
The postgres documentation will tell you that you'll need to set up your own mechanisms for this, and that they will need to integrate with OS facilities as appropriate. One-size-fits-all does not cut it. Not wrt. replication, not wrt. HA/failover.
All the tooling that provides extra distributed functionality not present in postgres (auto failover, multi master replication, sharding etc) will surely have issues, but then you aren't testing the PostgreSQL itself, but the tooling, so to be fair, you the article should evaluate these tools, and any shortcomings shouldn't go to PostgreSQL (unless it really is a PostgreSQL issue).
And looking at this table, basically the future seems to be WAL shipping anyway ( https://www.postgresql.org/docs/13/different-replication-sol... )
However, if I may suggest, Stolon, Patroni, Postgres XL or Citus Data might be interesting to you.
Even common highly available configurations take the route of no consistency guarantees by doing primitive async replication and primitive failover.
In a classic single node configuration, a confirmation that its transaction isolation behaviors exhibited the corresponding anomalies would be valuable.
So I think there’s value in this ask.
I think he did something similar for MySQL when evaluating the Galera cluster.
In a single write master configuration, Postgres runs transactions concurrently, so the consistency analysis is still quite relevant.
I don’t think it’s a stretch to say that everyone expects Postgres to get top marks in this configuration and it would be worth confirming that this is the case.
But it was long ago, and maybe needs to be redone?
Edit: after re-reading it he treats it as a distributed system because client and server is over network. And that is true, it can also be thought of as a distributed system because as you said transactions are concurrent and are running as separate processes. Although in these cases you can't have a partition (which aphyr uses to find weaknesses), or maybe there is something equivalent that happens?
Not in itself, but it does offer a PREPARE TRANSACTION - COMMIT PREPARED / ROLLBACK PREPARED extension that could be used to add such support in the future. This would not be unprecedented, as the simpler case of db sharding is already being supported via the PARTITION BY feature, combined with "FOREIGN" database access.
Very happy for (informed) speculation here, I recognise we'll probably never know for certain, but I'm interested to avoid making similar mistakes myself.
The middle part of the report talks about unexpected but (almost all) documented behavior around read and write concern for transactions. I don't want to conjecture too much about motivations here, but based on my professional experience with a few dozen databases, and surveys of colleagues, I termed it "surprising". The fact that there's explicit documentation for what I'd consider Counterintuitive API Design suggests that this is something MongoDB engineers considered, and possibly debated, internally.
The final part of the report talks about what I'm pretty sure are bugs. I'm strongly suspicious of the retry mechanism: it's possible that an idempotency token doesn't exist, isn't properly used, or that MongoDB's client or server layers are improperly interpreting an indeterminate failure as a determinate one. It seems possible that all 4 phenomena we observed stem from the retry mechanism, but as discussed in the report, it's not entirely clear that's the case.
I get the impression that MongoDB may have hyped themselves into a corner in the early days with poorly made (or misleading) benchmarks. Perhaps they have customers with a lot of influence determining how they think about performance vs consistency.
Maybe this combined with patching, re-patching, re-patching again their replication logic/consistency algorithm means that they'll be stuck in this sort of position for a long time.
This section seems to be the most worrying results in your report, Kyle, with no work around. Did I read that correctly?
That's not to say that workarounds don't exist, just that I didn't find any in the documentation or by twiddling config flags in the ~2 weeks I was working on this report. :)
So I'm curios how would you have described the ability of finding violations with Elle using read-write registers with unique values vs the append-only lists?
If you look at Elle's transaction generators, you can cap the size of any individual key, and use an uneven (e.g. exponential) distribution of key choices to get various frequencies. That way keys stay reasonably small (I use 1-10K writes/key), some keys are updated frequently to catch race conditions, and others last hundreds of seconds to catch long-lasting errors.
So I'm curios how would you have described the ability of finding violations with Elle using read-write registers with unique values vs the append-only lists?
RW registers are significantly weaker, though I don't know how to quantify the difference. I've still caught errors with registers, but the grounds for inferring anomalies are a.) less powerful and b.) can only be applied in certain circumstances--we talk about some of these details in the paper.
https://web.archive.org/web/20150312112556/http://blog.found...
Jepsen draws inspiration from a long line of work on property-based testing, especially Quickcheck & co. It also draws on roughly 10 years of experience building & running distributed systems in production. A lot of Jepsen I invented from whole cloth, but some of the checkers in Jepsen are derived from specific research papers, like work by Wing, Gong, and Howe on linearizability checking.
Then it's just... a lot of thinking, experimenting, and writing. Jepsen's the product of ~6 years of full-time work. Elle, the system which detected the anomalies in this report, was a research project I've been puzzling over for roughly two years.
I write the Jepsen series, and open-source all of the code for these tests, partly as a resource so that other people can learn to do this same kind of work. :-)
My bias: I like and heavily use ZooKeeper in production. HN seems not to like it as much.
For as easy as it is to use jsonb in Postgres, or Redis, or RocksDB/SQLite, or whatever else depending on your use case - I can't find any reason to advocate its use these days. In my anecdotal experience, the success stories never happen, and nearly developer I know has an unpleasant experience they can share.
Big thanks to aphyr and the Jepsen suite (and unrelated blog posts like Hexing the Interview) for inspiring me to do thorough engineering.
Anyone have any suggestions for a true non-MongoDB jsonDocument based noSql option?
Because the answer is "no" in the overwhelmingly majority of cases, specially if your product is mature.
Given that is a pretty popular part of MongoDB seems like an important thing for people to continuously fail to mention.
Your attitude of "a tool I need doesn't exists, so I'll just go ahead and create it" blew my mind and changed me for the better.
I'm dedicating my next test framework to you. Thank you for everything.
As an engineer for whom automated testing tools are crucial to my mental health, let me know if you want a UX tester or just someone to provide feedback on the documentation.
Sometimes you end up with bad defaults simply by accident but I feel like for MongoDB the morally correct choice would be to own up to past mistakes and change the defaults rather than maintain a dangerous status quo for "backwards compatibility", even if you end up looking worse in benchmarks as a result.
A detailed description of the bug can be found here: https://jira.mongodb.org/browse/SERVER-48307
This bug has been fixed and backported, and will be available to users in MongoDB 4.2.8 onwards. The MongoDB test suite has been updated to ensure that this specific phenomenon is detected in future releases. We are also planning to update the version of Jepsen we are currently running in our CI loop to include the newest test case used in the report.
Last, we’ve made some changes to how we share information discussed in the Jepsen reports on our website. You can find the updated page here (https://www.mongodb.com/jepsen).
I have been responsible for <100 clustered Cassandra instances, and <500 clustered MongoDB instances, and I would choose the latter every time.
I hope they learned the lesson, don't fuck with aphyr.
I’d love to read a roasting like that authored by Leslie Lamport for a different perspective but aphyr’s works absolutely stand on their own.
Any ideas how to get Jepsen and TLA to work together? :)
Their C/C++ client is literally unusable. I went to look into writing my own that actually worked and their network protocols are almost impossible to understand. BSON is a wreck and basically the whole thing discouraged me from ever trying to interact with that project again.
[0]: https://news.ycombinator.com/item?id=6801970 (BTW: no, my dream of simple migration never materialized, but exporting and dumping data to Postgres JSONB columns and rewriting queries turned out to be neither buggy nor hard).
This report is 9 days old, and tests the latest stable release of MongoDB. The problems it discusses are present on modern MongoDB.
I'm glad now that it's been clarified :)
I'm also excited about my own research with Elle, but we're still working on getting that through peer review, haha. ;-)
Woah, that's wild. Are there any pre-prints/papers/talks that you can link to on this subject? I'd _love_ to read this.
> I'm also excited about my own research with Elle, but we're still working on getting that through peer review, haha. ;-)
I read over bits of Elle; the documentation in it is absolutely top-notch. You and Peter Alvaro knocked it out of the park!
This is based on her presentation and some dinner conversation at HPTS 2019, so I don't know if there's actually a paper I can point to. The gist of is that Paxos normally involves an arbitration phase where there are conflicting proposals, which adds a second pair of message delays. But if you relax the consensus problem to agreement on a set of proposals, rather than a single proposal, you don't need the arbitration phase. Instead of "who won", it becomes "everyone wins". Then you can impose an order on that set via, say, sorting, and iterate to get a replicated log.
I read over bits of Elle; the documentation in it is absolutely top-notch. You and Peter Alvaro knocked it out of the park!
Thank you! Could I... hang on, just let me grab reviewer #1 quickly, I'd like them to hear this. ;-)
This sounds very similar to atomic broadcast (https://en.wikipedia.org/wiki/Atomic_broadcast) where each node sends a single message and the process ensures that all nodes agree on the same set of messages. Not sure how it would fit with a log-oriented FSM, but it certainly sounds interesting.
https://os.zhdk.cloud.switch.ch/tind-tmp-epfl/394a62dd-278f-...
Thanks for the explanation! I just found http://www.hpts.ws/papers/2019/howard.pdf; I'm reading through it now :)
> Thank you! Could I... hang on, just let me grab reviewer #1 quickly, I'd like them to hear this. ;-)
Do as you please with my praise!
Since they’re claiming something provably false, it’d be nice to have some empirical evidence as such.
I wonder if someone can type up a well-manicured post-Morten of the recent triple byte incident?
I understood that reference
MongoDB's also published a writeup (which is cited a few times in the Jepsen report!) talking about the impact of stronger safety settings and why they choose weak defaults: http://www.vldb.org/pvldb/vol12/p2071-schultz.pdf
Compare and contrast with the highly ethical Postgres team, who encourage good practices from the start and who get a feature right first before worrying about performance. That may harm their adoption in the short term but over the long term, that's why they're the gold standard. And with their JSONB datatype they have a better MongoDB than MongoDB anyway! And have a million other features besides!
Yeah, but in spite of that their performance still sucks compared to writing directly to /dev/null, and that's where Mongo steals their thunder.
You do know that PostgreSQL had issues with not fsyncing data as well ? It's technology. Bugs will be made. Design decisions will be wrong.
I think it's really disappointing and inappropriate to be labelling MongoDB engineers as unethical for simply having incorrect defaults. Which in their history they often change after they are made aware of them.
See, you can name just one Postgres bug, and they held their hands up to it straight away. Whereas the MongoDB "bugs" are countless and by sheer coincidence, they mostly skew to improving performance in benchmarks and demos. That's a pattern.
There's no reason for Jepson to be applied to a single-node in-memory kv store.
Now that it's just "MongoDB 4.2.6", the title makes me think that this is a release announcement, not an analysis of the software.
The first title (that specifically referenced a finding of the analysis) was best, imo. Mildly opinionated or whatever, but at least it quickly communicated the gist of the post. On the other hand:
"Jepsen: MongoDB 4.2.6" – not super helpful if you're not already familiar with the Jepsen body of work.
"MongoDB 4.2.6" – as stated above, sounds like a release announcement.
If you want a suggestion, maybe something like "Jepsen evaluation of MongoDB 4.2.6"? Not overly specific (/ negative) like the first title, but at least provides some slight amount of context.
@dang
I don't mind making an exception, since exceptions are things sometimes. Jepsen is famous on HN, so the current title is not an issue. Indeed, referencing a specific finding would arguably be misleading, since this article is the Jepsen report about MongoDB 4.2.6. Btw, I don't know what you mean by "The first title (that specifically referenced a finding of the analysis) was best". The submitted title was "Jepsen: MongoDB 4.2.6" and it has only ever rotated between two states, one with "Jepsen: " and one without. Are you confusing this thread with https://news.ycombinator.com/item?id=23285249?
It's very silly to have this be the top comment on the page (I've since downweighted it, but that's where it was when I looked in). Yesterday I briefly swapped the URL of this article into the other thread, but then reversed that because it seemed that thread couldn't support a more technical discussion (https://news.ycombinator.com/item?id=23288120). I invited aphyr to repost it instead, which was quite a break from our standard practice of downweighting follow-up posts, but seemed like the best solution at the time. What technical discussion was our reward? Bickering about title policy!
"Why don't you put Jepsen:" on the same line as the database name and version?"
Space concerns, and also, it's immediately above the DB name in giant letters.
"Why don't you give them more creative names?"
Clients love to argue about the titles of these analyses; having a concise, predictable policy for titling is how I get past those discussions.
But then again ultimately the blame is on the author of the article, it's a terrible title for this type of articles. I can understand if the moderators here don't want to go through the trouble of dealing with editorialized titles (with all the controversies it could generate) when clearly the original author didn't care enough to come up with a decent title.
Site context isn't a given when most of us are finding content via 3rd party sources.
So you manually moderate the content?
I, for one, welcome this by-hand moderation because it keeps this issue alive, and allows Kyle to keep the discussion going.
As I commented in a previous post, Kyle is the Chef Ramsey of database testing, and here, he's in a position where some idiot has just served him an undercooked hamburger. Bits will fly, marketing people will be flayed alive, and Kyle will be the only one left standing at the end.
Without this by-hand moderation, we'd be missing out on the second act of this intense thriller!
https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
(I've detached this subthread from https://news.ycombinator.com/item?id=23294048 to prevent the top comment from being too distracting.)
It's totally fine with me, but I just wasn't aware of it.
No clue if this "downweighting" in this case is an algorithm or a manual thing. I would assume algorithm for the downweighting and human intervention for reversing it, but that's sort of a guess or inference.
Majority write/read concern is exactly so that you don't loose data and don't observe stuff that is going to be rolled back. It is important to understand this fact when you evaluate MongoDB for your solution. That it comes with additional downsides is hardly a surprise, otherwise there would be no reason to specify anything else than majority.
You just can't test lower levels of guarantees and then complain you did not get what higher levels of guarantees were designed to provide.
It is also obvious, when you use majority concern, that some of the nodes may accept the write but then have to roll back when the majority cannot acknowledge the write. It is obvious this may cause some of the writes to fail that would succeed should the write concern be configured to not require majority acknowledgment.
The article simply misses the mark by trying to create sensation where there is none to be found.
The MongoDB documentation explains the architecture and guarantees provided by MongoDB enough so that you should be able to understand various read/write concerns and that anything below majority does not guarantee much. This is a tradeoff which you are allowed to make provided you understand the consequences.
Gently, may I suggest that you read the report, or at least the abstract? This is addressed in the second sentence. :-)
Of course, it doesn't work when you don't pair it with majority read/write concern. You can't expect to get a snapshot of data that wasn't yet acknowledged by majority of the cluster.
As to the quote you probably are referring to:
"Jepsen evaluated MongoDB version 4.2.6, and found that even at the strongest levels of read and write concern, it failed to preserve snapshot isolation."
I did not find any proof of this in the rest of the report. It seems this is mostly complaint of what happens when you mix different read and write concerns.
I would also suggest to think a little bit on the concept of snapshot in the context of distributed system. It is not possible to have the same kind of snapshot that you would get with a single-node application with the architecture of MongoDB. MongoDB is a distributed system where you will get different results depending on which node you are asking.
The only way you could get close to having a global snapshot is if all nodes agreed on a single truth (for example single log file, block chain, etc.) which would preclude read/write with concern level less than majority.
"Tansactions without an explicit read concern downgrade any requested read concern at the database or collection level to a default level of local, which offers “no guarantee that the data has been written to a majority of replicas (i.e. may be rolled back).”"
The big problem is that, even if somebody correctly sets the read and write concerns to something sensible, the moment they use a transaction these guarantees fly out the window, unless they read the docs carefully enough to realise they have to set the read and write concern for the transaction too. The defaults are very un-intuitive; I can't imagine that the case of somebody needing snapshot isolation in general but being fine with arbitrary data less in transactions is a common case, compared to wanting to avoid data loss both generally and in transactions.
Yes it works. Yes, you have to read the documentation very carefully.
> transactions running with the strongest isolation levels can exhibit G1c: cyclic information flow.
As well as the Node.js API issue (I just checked randomly and their Python API has the same bug lol) listed above.
If the Stripe API had documentation was needlessly unclear in a way which led people to lose a significant amount of money, that would be a bug.
May I suggest sections 3.4, 3.5, 3.6, 3.7, 4.0, and 4.1?
This anomaly occurred even with read concern snapshot and write concern majority
3.5: In this case, a test running with read concern snapshot and write concern majority executed a trio of transactions with the following dependency graph
3.6: Worse yet, transactions running with the strongest isolation levels can exhibit G1c: cyclic information flow.
3.7: It’s even possible for a single transaction to observe its own future effects. In this test run, four transactions, all executed at read concern snapshot and write concern majority, append 1, 2, 3, and 4 to key 586—but the transaction which wrote 1 observed [1 2 3 4] before it appended 1.
Like... if you had read any of these sections--or even their very first sentences--you wouldn't be in this position. They're also summarized both in the abstract and discussion sections, in case you skipped the results.
4.0: Finally, even with the strongest levels of read and write concern for both single-document and transactional operations, we observed cases of G-single (read skew), G1c (cyclic information flow), duplicated writes, and a sort of retrocausal internal consistency anomaly: within a single transaction, reads could observe that transaction’s own writes from the future. MongoDB appears to allow transactions to both observe and not observe prior transactions, and to observe one another’s writes. A single write could be applied multiple times, suggesting an error in MongoDB’s automatic retry mechanism. All of these behaviors are incompatible with MongoDB’s claims of snapshot isolation.
It's OK to stop digging now!
Compared to a product like Oracle, transactions on MongoDB are very new, very niche functionality. Even MongoDB consultants do openly suggest not to use it.
MongoDB is really meant to store and retrieve documents. That's where the majority read/write concern guarantees come from.
As long as you are storing and retrieving documents you are pretty safe functionality.
Your article presents the situation as if MongoDB did not work correctly at all. That is simply not true, the most you can say is that a single (niche) feature doesn't work.
Have you ever tried distributed transactions with relational databases? Everybody knows these exist but nobody with sound mind would ever architect their application to rely on it.
Any person with a bit of experience will understand that things don't come free and some things are just too good to be true. MongoDB marketing may be a bit trigger happy with their advertisements but it does not mean the product is unusable, they just probably promised bit too much.
I am delighted to say that yes: checking safety properties of distributed systems, including those of relational databases, is literally my job. See https://jepsen.io/analyses for a comprehensive list of prior work, or http://jepsen.io/analyses/tidb-2.1.7, http://jepsen.io/analyses/yugabyte-db-1.1.9, http://jepsen.io/analyses/yugabyte-db-1.3.1, or http://jepsen.io/analyses/voltdb-6-3 for recent examples of Jepsen analyses on relational databases.
Holy shit, buddy. Stop.
So don't worry about me.
What I do have an interest in is HN's accepted decorum, which I admittedly stepped outside of when I implored you to stop digging yourself such a hole.
HN is far from perfect but there is a culture of respectful discourse here, which is part of the reason for its value IMO.
Can't reply to that since it's too nested so I'll reply here. I warmly recommend getting off tree you climbed on and actually reading the article because if you do - you will see you are not disagreeing on that part.
The article is a mostly technical analysis of the transaction isolation levels and where they hold. The main criticism is how MongoDB advertises itself. If they didn't claim the database is "fully ACID" then the article would have just been a technical analysis :]
As someone who is a tech lead for a large database install, I'd urge you to read the rest of the Jepsen reports. They aren't intended to be hit pieces on technology - they're deep dives into the claims and guarantees of each database. IIRC MDB has explicitly reached out to OP in the past (I doubt they'll continue to do so after this).
Why that matters to the rest of us: once I learn all those dials and knobs I'm left wondering why I would choose Mongo over another technology, and how much the design of the default behavior and complexity of said dials/knobs are influenced by their core business.
Imagine there was a programming language which had rather inconsistent naming, poor automated testing support, and a history of guiding its users toward security vulnerabilities. A culture would grow up around that language and the most successful members would be those who could best tolerate those properties. People generally self-select into language communities. So unless some powerful influence pushed random programmers to use the language or made it easier to add new tooling, the culture would continue to undervalue what the language originally lacked.
I suspect the same social dynamic would apply to a database.
If not, that's misrepresentation.
In any case any person that has some experience with distributed systems will understand what it roughly means to get an acknowledgment from just a single node vs. waiting for the majority.
Oracle also does not use serializable as its default isolation level, yet it advertises it.
This is all part of the product functionality. Whenever you evaluate product for your project you have to understand various options, functionalities and their tradeoffs.
Defaults don't mean shit. In a complex clustered product you need to understand all important knobs to decide the correct settings and configurable guarantees are most important knobs there are.
Thanks, Those of us who care about our banking and investing data.