A Gentle Introduction to CRDTs
vlcn.io
vlcn.io
D1, unfortunately, doesn't support loading extensions into SQLite.
fly.io, and their sqlite solution, does however!
I've been using fly.io to run cr-sqlite in the cloud and act as a sync server for non peer to peer use cases. Eventually I'll have finished up a re-write of strut.io[1] atop vlcn/cr-sqlite which'll serve as a good end to end example of how to make applications with this sort of architecture.
Some discussion about this setup on the fly.io forums: https://community.fly.io/t/litefs-many-tens-to-hundreds-of-t...
It's basically an offline-first flashcard webapp. CR-Sqlite allows for incremental syncing.
With Anki (the app from which I'm taking my inspiration), syncing is _not_ incremental - basically it just copies SQLite files around. So for example, the app could be on an iPhone with cards a card `A` reviewed, but the app on an iPad could make changes to the template on which card `A` is based, and that's enough to cause a conflict - you must take changes from only the iPad or only the iPhone. (To be clear - Anki does have some incremental syncing capabilities - I'm picking an intentionally pathological example.) CR-SQLite will mean that everything is incremental, however.
Basically makes 3 way merges a breeze (or n-way merges, really).
Wonder if/how it would combine with the distributed litestream stuff recently touted on fly.io? I love postgres but have to admit these features are tempting.
It's kind of a mindset shift, but with logical clocks you kind have to let go of the concept of ordering with one global universal time - it's just not feasible to do accurately with clock skew.
They're purely used to provide a partial order of events - partial because the system cannot say for sure whether some events happened before or after each other.
Even more fundamentally, it’s not possible because of physics. The passage of time is relative!
Not as important for terrestrial applications, but useful to consider as a theoretical limitation.
But you're bang on the money with the physics analogy. Lamport himself says that his knowledge of special relativity helped him come up with the concept.
This seems to be such a nice construction, but I don't know its name.
Just like any other physical measures. You don't measure nearly any of them with a wristwatch (only proper time/aging) but with a device that is constructed according to the definition of that physical measure.
This said, in a truly decentralized setting, where CRDTs make most sense, logical clocks are not useful as they are not not byzantine fault tolerant. You need merkle trees.
byzantine fault tolerance is fun and all but it's a niche use case
You can only trust the nodes that you control. That's not what I'd call a decentralized system.
If we see every field as f(x,y,z,t) what do you mean you can't "change" x,y,z without changing t?
If you mean changing the x,y,z argument values, you can (you're just evaluating a different point in space).
If you mean changing the evaluated function value at a given x,y,z it's kinda tautological that it stays constant for a constant x,y,z,t.
So when time stops, everything else stops. It is tautological when you think about it, but I find that it's an easy way to explain why logical clocks work.
Am I wrong to reason this way?
Imagine there is an object recorded at location x,y,z at time t. This is all we know.
If I later tell you the object was recorded at location x,y,z, you have no idea what time it was captured. If I later tell you the object was recorded at any location other than x,y,z (even if only z changed by some epsilon) you know that that it was captured at some time other than t.
If you guarantee the second recording I'm showing you was not taken before time t, you know it is a more recent recording, because it would be not earlier than and not equal to time t.
Other GC research unrelated to this project but something I've been following: https://braid.org/antimatter
I'll try to get to that.
For a simple case of two nodes, they'll be divergent until each node sends the other their current state.
For client-server, the "whole system" wouldn't be in the same state until all clients report all their changes to the server and the server re-broadcasts all those changes to all clients.
A pathological case (but happens when people leave an ipad in a drawer or something) -- a node could be offline for 6 months and the rest of the system never gets those changes until that node comes back online. There's technically some divergence: the other nodes are missing whatever state the ipad had, but does it really matter if that node was unimportant enough to be left offline?
Depends on use case for sure.