While quorum looks a lot like consensus, it is not since what is returned to the client is the latest timestamp. Different nodes could have and return different data. So Quorum write+local_quorum read might fail even if you are in the datacenter that accepted the write. Quorum is also on the total copies of the data, not a quorum of DC, so in certain (weird) multi-DC setups you could have a quorum in a single DC.
In general though I think that the data consistency options of cassandra (quorum/local/N) are a good idea but underdeveloped.
Anyway my point was: cassandra has too many pitfalls and eliminating those restricts the use case by a lot more than people realize. Plus the naming of all features look designed to trick you into thinking soewthing else
Yes, if you mix consistency levels incorrectly you will do it wrong and maybe get stale data, but that is a different criticism. I agree that it is easy for unsophisticated users to incorrectly use consistency levels in complex topologies, and I hope we will introduce mechanisms to prevent users making such mistakes in future. But that was not your claim, and in my experience users do understand consistency levels just fine.
There are lots of valid criticisms to point at various use cases with Cassandra, but this was just incorrect.
If value 2 is newer than value 1, value 1 will never overwrite value 2. If a client reads value 2 at QUORUM then it will always be seen by all future queries.
then later client B write value 2 on 2 of the 3 node with timestamp=1
then read repair happen and (value=1 timestamp=3) is written to all 3 nodes.
this is only one of many scenarios where this stupid design fail.
This would have been a valid criticism of LWW (and there are other more contrived examples), but I think (or hope) this is an explicit trade off made by anyone using Cassandra in eventual consistency mode. There are strategies to prevent this being a problem for workloads where it matters, some discussed elsewhere in the thread.
Quorum Read-repair is only one reason for why the value would randomly disappear. Another one is periodic anti Entropy repair!
Since at least one of those nodes has the “newer” value, only one node can serve this “older” value
Only if they miss some writes, and eventually they will converge. But if you do quorum writes and quorum reads (or local quorum W + local quorum R), this guarantees you'll read from at last one node that received all the writes issued before the read, so you get the converged value immediately, regardless of which node you ask. All nodes will eventually agree on the value, because timestamps are assigned by the coordinator or by the application, not at the replica. A single write will get the same timestamp across all replicas.
Incorrect timestamps can cause a different problem - a write that happened at physical time T2 > T1 might be considered to be older than T1 by the cluster, if it was accepted by the coordinator whose clock was set in the past. Such write might simply not take any effect, as old updates would be considered newer. However, again, the resolution would be consistent on all replicas once they get all updates.
you are guaranteed to face the problem where the client machine clocks are not in sync.
just comparing timestamp is obviously a design from someone that didn’t review the academic literature on distributed transaction and consensus
The client can set the timestamp to any value of its choice. It does not have to correspond to clock time.
>> so unless all your write request are issued by the same machine then you are guaranteed to face the problem where the client machine clocks are not in sync.
It's not about all writes. It's about ensuring that all writes to a single partition (specifically, a single column within a single partition) are done using a source of monotonic integers to ensure ordering.
But this Monotonic time server is not part of Cassandra itself and for this reason majority of the people using Cassandra will use the OS time without knowing this is silently corrupting the database.
This is all done for you by Cassandra
If you have 3 nodes, set quorum cl, the write will succeed, because quorum of 3 is at least 2.
If you explicitly require cl of three, and only two nodes answer, write will fail.
Unless the node was down too much and could not fully catch up before a set time (DB TTL if I remember correctly), in which case the data might be propagated or not, and old deleted data might come up again and be repropagated by a cluster repair.
So much fun to maintain
That pretty much describes iCloud.
Argh! Zombies!
iCloud has terrible syncing. Here’s an example. Synced bookmarks in Safari:
I use three devices, regularly; my laptop (really a desktop, most of the time), my iPad (I’m on it, now), and my iPhone.
On any one of these devices, I may choose to “favorite” a page, and add it to a fairly extensive hierarchy of bookmarks, that I prefer to keep in a Bookmarks Bar.
Each folder can have a lot of bookmarks. Most are ones that I hardly ever need to use, and I generally access them via a search.
I like to keep the “active” ones at the top of the folder. This is especially important for my iPhone, which has an extremely limited screen (it’s an iPhone 13 Mini -Alas, poor Mini. I knew him well).
The damn bookmarks keep changing order. If I drag one to the top of the menu, I do that, because it’s the most important one, and I don’t want to scroll the screen.
The issue is that the order of bookmarks changes, between devices. In fact, I just noticed that a bookmark that I dragged to the top of one of my folders, yesterday, is now back down, several notches.
Don’t get me started on deleting contacts, or syncing Messages.
I assume that this is a symptom of DB dysfunction.
Otherwise it's far more likely to be faulty sync algorithm.
If you are selecting at ONE then yes, you can expect stale replies if you contact a different node. That is what the consistency level means; that just one node has seen the write.
The other behaviour you’re talking about is hinted handoff.
“So much fun to maintain” -> cassandra has very specific use cases. And hosted versions exist.
Disclosure: I work at Aiven, which has an hosted cassandra offer.
https://cassandra.apache.org/_/blog/The-Path-to-Green-CI.htm...
https://cassandra.apache.org/_/blog/Finding-Bugs-in-Cassandr...
https://cassandra.apache.org/_/blog/Testing-Apache-Cassandra...
https://cassandra.apache.org/_/blog/Introducing-Apache-Cassa...
it mostly work but it’s extremely slow and still have many bugs.
https://www.datastax.com/blog/lightweight-transactions-cassa...
Until recently they were indeed slow over the WAN, and they remain slow under heavy contention. They are now faster than peer features for WAN reads.
However, the claim that they have many bugs needs to be backed up. I just finished overhauling Paxos in Cassandra and it is now one of the most thoroughly tested distributed consensus implementations around.
LWT still have awful performance compared to a write request with same guarantee in any other database.
https://thenewstack.io/an-apache-cassandra-breakthrough-acid...
A prototype has been developed that has demonstrated its correctness against Jepsen.io’s Maelstrom tool