A Certain Tendency of the Database Community
christophermeiklejohn.com
christophermeiklejohn.com
Not all updates are commutative.
The ordering of events in (add 1, add 2, add 3) does not affect the outcome.
The order of events in (add 1, multiply 2, divide 3) does matter.
The problem of linearity exists as soon as you have a commutative operation in a relativistic universe where information cannot be simultaneously updated in two coordinates simultaneously.
Until now there have been heroic engineering efforts to hide this from software developers. Hardware has conspired to conceal that even in a space a few centimetres wide, values can become inconsistent.
Once you have independent memory spaces, all bets are off and you are forced to pay a high price for losing the abstraction.
But on my reading, CRDTs just kick the can down the road to the implementer of the operation or merge function. The way they solve for commutativity is to ... require commutativity.
That's a win if your changes are commutative. Lots of things aren't.
Any and all arithmetic operations can be fully solved with CRDTs without further can kicking. The trivial solution is to give every number you are adding/subtracting a unique identifier, then as the values come in you do a sum across them, deduplicating in the process. This means you can resend the math operations as many times as you want and you still get an idempotent result.
To get a consistent answer, you'd also have to sort the numbers consistently, so instead of a sum this becomes an insert into a sortable list. This is far more expensive than the typical kinds of aggregate operations people want to do at scale - the point of aggregation is not to just save all the original data.
How would this system work when treating untrusted clients executing arbitrary code as the definite source of truth for a fact? Further, it doesn't seem like this addresses (unless I'm missing something) the problem raised in the airline ticket example where split-brain systems may hand out resource reservations that they aren't actually allowed to give out. If both passengers' devices (or the plane? or some other edge device?) are sources of truth, I don't see any easier way to negotiate the resource contention than current distributed databases already have.
However, in the real world, waiting for consensus before making decisions is hard, so airline systems, designed in the era before realtime communication back to the mothership was possible, faked it. They made a reservation, issued a ticket, and then the airline developed protocols to deal with the resulting overbooking. In fact, since airlines already had incentive to overbook and deal with overbooking because of no-shows, they simply built the failures of their locking system into their cost-models.
There is a 'correct' solution to this: When you believe you're going to be split-brained some of the time, pre-allocate tickets to entities. Those entities can use their local tickets, but cannot make use of remote tickets without approving them with the systems that own them. Sharding the state of the plane down to individual seats has system complexity cost, but allows perfect booking.
And even if you design the system to operate in your 'correct' solution, you may still result in imperfect booking if one side of the split exhausts its pre-allocated tickets and reports as unavailable when other sections of the system still have available tickets.
There are large numbers of problems that are easily amenable to casual consistency, but there are other that simply are not. The possible solutions to the former are many and varied but the latter requires a solution far more in tune with the "database community."
I think people need to know that this discussion is no longer happening only in the hands of academics, this is now available to developers - with or without the academic side.
It doesn't look like you have. I watched the demo video "GUN Survives a Primary Fault" and the conflicting updates were resolved with a LWW semantic, which is non-deterministic.
(EDIT: I read some docs, and it looks conflicts are resolved by choosing the highest lexical order of the value. So, apologies for the assumption. I assumed LWW because the BBBBBB update was seen last after the fault was resolved. Nonetheless, I think it is interesting you're treating the browser as a peer in the system, and not simply a client)
> this is now available to developers - with or without the academic side.
I don't understand this attitude. Working closely with academics ought to be what industry strives for. They take the time to research, formalize theorems and write proofs to advance the current state of the art in Computer Science. This research helps move industry forward. A tight relationship between industry and academia is mutually beneficial, for what should be obvious reasons.
Academia - I agree with you! Good words, sir. I was just wanting people to know that it does exist in code land as well, because often times academic designs aren't. So I think it is important to highlight when they are available.