The longer you're offline the more you have to explain to your users the circumstances where they might need to do manual data cleanup or expect creative merging policy. No matter what tech you're using.
The longer you're offline the more you have to explain to your users the circumstances where they might need to do manual data cleanup or expect creative merging policy. No matter what tech you're using.
How SSB works is a good example.
But yes, it’s possible there are applications that can constrain the set of operations such that the merged result is also correct. I was just highlighting that CRDTs are trotted out as a catch all when a) it’s just a way of thinking about the problem rather than an actual algorithm (eg while there’s a lot of CRDT work for document editing, that work has to be done for scratch for any other problem domain) b) you have to modify your problem constrains and solution such that merge results are correct. B is a very very hard problem in a distributed system even if you ignore the challenge of a which itself is quite hard. I wouldn’t trust anything in this space too much without a TLA+ proof that it’s correct (unless it’s something low stakes like document editing).
Two users simultaneously select the same word, and replace it by a different word each.
You might end up with: "apple" -> "carrot orange"
Sure it technically merged, and did not lose data. But you still need humans to reconcile the final state anyways. Worse, you might not even notice that you need to cleanup.
The easier scenario that is independent of the specific algorithm is the one I described (have follow up comments describing it in more explicit detail) where you have a logical meta instruction within the document and then the document is updated based on outdated meta instructions offline but when they come back online the meta instructions were changed.
Then there’s not even a merge conflict to really worry about.
The CRDT may pick one or the other replacement word, but who is to say that either choice is correct? Perhaps including both words is correct.
> Then there’s not even a merge conflict...
Agree, this is what CRDTs are all about.
> ...to really worry about.
I think it is important to make clear that CRDTs do not "solve" the merging problem, they merely make it possible to solve in a deterministic way across replicas.
Often, CRDTs do not capture higher level schema invariants, and so a "conflict free" CRDT merge can produce an invalid state for a particular application.
There is also the example above, where at the application level, one particular merge outcome may be preferred over another.
So, it isn't as simple as having nothing to worry about. When using CRDTs, often, there are some pretty subtle things that must be worried about. :-)
You, as the application developer.
> I think it is important to make clear that CRDTs do not "solve" the merging problem
They literally do, in the context in which they are defined. Which is about data consistency, not semantic correctness.
> Often, CRDTs do not capture higher level schema invariants,
CRDTs never capture higher level schema invariants. Just like TCP doesn't enforce HTTP session authentication. Orthogonal concerns.
It's not technically a conflict since there is a path forward, but the resulting output can be nonsensical.
For example you could employ a strategy where given two concurrent edits, the merge "randomly*" picks one of the two edits for each property/fact, and abandons the other.
That's a construction of a conflict free replicated data type, but it says nothing of the quality of the final result.
Ultimately, the quality comes down to how well you can produce merge strategies for your domain.
* "Randomly" in a deterministic way, e.g. by comparing a hashes of transaction IDs or similar.
...which is the sense that we're talking about, when we talk about CRDTs.
This is a bonkers conversation. It's like claiming integer division isn't correct because 5/2 should be 2.5 instead of 2.