RavenDB 6.0.2 (A Jepsen Report)
jepsen.io
jepsen.io
Interesting trivia: there's "raven db done right" - https://martendb.io/ , just an API wrapper around PSQL. Named Marten because thats a natural enemy of ravens:)
MartenDb is great and the community around it is excellent.
- Replication is a technique used to achieve higher availability and durability than a single node can offer, by making multiple copies of the data. Techniques include Paxos, Raft, chain replication, quorum protocols, etc.
- Active-active means that transactions can run against multiple different replicas at the same time, while still achieving the desired level of isolation and consistency.
- Atomic commitment is a technique used in sharded/partitioned databases (which themselves exist to scale throughput or size beyond the capabilities of a single machine) to allow transactions to be atomically (“all or nothing”) committed across multiple shards (and allow one or more shards to vote “nah, let’s not commit this”). 2 phase commit (2PC) is the classic technique.
- Concurrency control is a set of techniques to implement isolation, which is needed in any database that allows concurrent sessions (single node or multi-node). Classic techniques include 2PL and OCC, but many exist.
When vendors or projects answer concurrency control questions with replication answers (which appears to be the case here), it’s worth diving deeper into those answers. There are cases where “Paxos” or “Raft” might be answers to atomic commitment or even concurrency control questions, but at best they are very partial answers and building blocks of a larger protocol. Databases that only support “single shot”/predeclared transactions can get away without a lot of concurrency control, for example, and might be able to do the required work as part of their state machine replication protocol. In general, I'd see using words like "Paxos" and "Raft" in the marketing for a database as a negative sign. It's not a fully reliable one, but it's often the least interesting part of the implementation and the choices the database is making.
To be extra clear, I’m not criticizing Aphyr here (the article clearly doesn’t conflate these concepts), but more pointing out what I think lies at the bottom of a lot of the issues we see with distributed database claims.
> I am not sure if talking about anything else in the context of distributed databases is reasonable.
There's a whole world in distributed databases, and I suspect you'd agree that there's a lot of stuff worth talking about that isn't covered in your (excellent) work.
You can’t force people to test their software correctly unless it’s a regulated field like aviation.
Either way, best to assume most companies are lying through their teeth about any feature until you or someone you trust has validated it.
> RavenDB’s Jepsen test may not have measured anything at all: at least in the most recent revision, the generator included no client operations of any kind.
Remember folks, if you can't get your test to fail by intentionally breaking the implementation, you don't have a test.
It did not... go exactly as planned. Initial tests looked OK, but when I did testing with actual users, there were huge issues right away. Like: OK, I just ran your ingestion pipeline. What do you see? And the answer was 'well, nothing', or 'ehhm, a lot less than I expected'. These issues turned out to be pretty much impossible to fix: there were no real errors, but the data just seemed to... disappear randomly, even in a simple single-node cluster. I got community support involved in a bunch of particular issues, but nothing really helped: the aggregate numbers we got never added up to what they should be.
I then migrated the whole thing to a single SQLite database. That file is, as I write this, a good 2TB in size, and still performs as well as the day it was deployed and never had any unexplained-number issues, without any changes to the surrounding code. I did eventually move away from the .NET Entity Framework (as that did cause some rare, yet unexpected and hard-to-fix concurrency issues, but those were hard crashes and not silent data corruptions) to a hand-rolled entity mapper, but all has been good since then...
TL;DR: databases are very hard, and fashionable choices are not necessarily desirable.
Won't go into battle scars here, but this report does not surprise me. We're much happier with Postgres and Elasticsearch.
I think in 99.9% of cases, you don't want AP. The P only matters when the network is more prone to go down than the machines. For example, if every node goes down, your AP design won't be available.
With the massive improvements in network and connectivity and increased redundancy, you should aim for CP.
If you really, really need AP, then a ground up design based on CRDTs seems the best, most discipline approach. With CRDT, you can have availability because the operations can be entirely local, and you know you can sync to the other nodes when available without conflict.
What are your most tempting/daunting databases that you haven't got a chance to put through their paces?
And, a second question, if you step back and think about the various APIs you've had to use, have you personally developed favourite styles of API to use?
I'd love to do more work with predicates in general. That's an open research problem I've been noodling on for years. Pretty much any SQL DB would be a good candidate for that work!
I'm gonna be a weirdo and say I actually loved Fauna's FQL. A little Lisp-ish functional language for queries is a great way to interact with document-structured data. SQL is fantastic for sheer breadth, though its specification is a nightmare and actually writing portable SQL is real challenging. One of those places where a stronger spec and conformance tests would have really helped.
If you allow me : how many people you work with can actually perform those Jepsen Report ? Or is it only you ?
There's a lot of folks out there who can do basic testing work with Jepsen. I've taught... I dunno, maybe a few hundred people directly in Jepsen workshops. A couple people have worked alongside me, and I'm sure lots more have learned from the docs online. Writing a report is a more involved problem--certainly not intractable, but for me it involves testing, experiment design, lots of reading, doc review, writing, editing, finding reviewers, and of course all the business stuff.
I don't know why all the DBMS vendors don't just have a guy on the QA team whose job is to run and interpret Jepsen tests for every new version. It's certainly a better option than eventually getting a damning report written by you.
As an anecdote i was surprised to discover mongodb had a second life in the corporate world as a standard , certified technology to store critical documents. So yeah, maybe people aren't really that aware of the kinds of nasty gotchas that lure in their systems.
But RavenDB does not do it justice and uses unsafe in catastrophic amounts in places where it is not necessary or in ways which are straight up UB despite the fact that JIT/ILC is much more strict than GCC/LLVM. There have been multiple bug reports submitted to dotnet/runtime by RavenDB which required extensive debugging effort only to end up being an issue on RavenDBs end due to explicit misuse of unsafe APIs (in ways, I must reiterate, that have safe alternatives to achieve the same performance).
(if anyone's interested, I can later ask around/dig through issue history and give the references)
I witnessed RethinkDB losing to MongoDB in spite of being significantly better. I am now worried that FoundationDB isn't gaining popularity, even though it is arguably the best and most well-tested distributed database out there, with strict serializability (!) guarantees. But it doesn't have a shiny website and doesn't cause warm fuzzies, quite the opposite, it looks complex and intimidating. So it isn't popular.
This is worrying, but perhaps neither new nor surprising: we have a history of picking inferior solutions because the good ones looked too complex or intimidating (betamax vs VHS in video formats, ATM vs Ethernet in WANs).
Can you share any more detail? Are you saying there are companies that build software on top of FoundationDB as their primary data store? or are those companies building software around FoundationDB that in turn presents as more of a data store in the traditional sense?
https://www.snowflake.com/blog/how-foundationdb-powers-snowf...
But yeah you're right for the most part. Turns out pretty much any database can be written in terms of transactions of KV pairs, which is what foundationdb gives you, so it means you can write your database query layer as a stateless, scalable service.
There have been attempts to write a SQL RDMS layer for it but it isn't maintained.
What you do not get is a "query language" or indexing.
Yes, there are companies that use FoundationDB as their primary data store. It makes a lot of sense to integrate it directly with the application rather than go through additional layers and a "query language". I am working on adapting my app to use it, and so far very happy with the results.
also, a lot more operational tooling would be nice.
I mean who posts that kinda stuff on their public website?
If you're not interested in a science experiment, Cassandra (or Scylla) are the multi-master databases that are mainstream and proven to work and scale. They're not fun or sexy and their feature set is much smaller but they do what they say and they work. Or AWS/GCP/Azure will happily give you an API for one.
That's different from how Raven has just wildly mistated their capabilities.
That database spent 3-4 years primarily focused on correctness from 2018-2022.
The industry moves fast, our memories are slow. But there are millions of instances of cassandra in production across most of the fortune 500, and half of this thread has never heard of RavenDB
The big difference is in how severe the difference between claims and reality are, and how the maintainers or vendors react to such bug reports.
In some cases, the maintainers or vendors will fix bugs, or update docs to be more clear. The different transaction isolated levels are complicated and nuanced, and there are standards which disagree with the general consensus in the literature, so there can easily be ambiguities that need to be cleared up or bugs that need to be fixed.
But then there are things like RavenDB; where they make clearly impossible claims like "ACID across multiple documents and multiple cluster nodes in an AP database." There is simply no way to achieve this.
And then there's how they respond to his findings. He filed bugs and had had responses a month ago, about things like "this thing that is documented to have transactional semantics does not have transactional semantics", and their response was just to say "oh, yeah, that's expected", and not fix anything or update any of their documentation to reflect that.
So, there's a big gulf between "this is a complex topic, and even some of the best systems have some issues if you test it thoroughly enough", and "the documentation is blatantly lying about transactions and consistency, and the CEO of the company doesn't think it's a problem."
The docs for RavenDB explicitly state https://ravendb.net/docs/article-page/6.0/csharp/client-api/...
The batched operations that are sent in the SaveChanges()
will complete transactionally. In other words, either all
changes are saved as a Single Atomic Transaction or none
of them are. So once SaveChanges returns successfully, it
is guaranteed that all changes are persisted to the database.
But the tests in this post show that no, there is no single atomic transaction for a session saved with SaveChanges, even in a single node database this will lose writes.I dunno. If I paid for an ACID database, I'd expect, well, some ACID features, like the ability to run two transactions concurrently and have them be isolated. It looks like RavenDB is fundamentally not implementing even the most basic of its claims. This isn't some "oh, yeah, this is a complex problem, and there are a few bugs lurking in obscure corner cases", this is "it fundamentally doesn't support what they claim to support as the major front-page selling point."
Their front page selling point is "Fully transactional NoSQL database" which links to a page that says "A database without transactions is… well, not much of a database, and as far as transactions are concerned – ACID is the gold standard."
But then in response to these findings, they say that the thing that is documented to have transactional semantics doesn't actually have transactional semantics.
Iike it actually more than SQL, especially for typical LOb apps.