You Might Not Need a CRDT: Document Sync in the Wild [video]
youtube.com
youtube.com
I work on Notion, a collaborate editor that (famously?) doesn’t merge text updates. Our editor predates the popular CRDT libraries. We’ve taken the last write wins optimistic model about as far as it will go, and now we’re building CRDT text as the basis for a bunch of new features. It’s challenging to integrate with our existing model but I think the end result will be well worth it.
If you’re interested in this sort of thing and want to work in SF or NY, shoot me an email jake@makenotion.com or dm me on Twitter @jitl
For an app like Notion you don't have direct peer-to-peer syncing: you have a star topology with a central service able to create a canonical total ordering of events, so you really don't need CRDTs at all.
Direct peer-to-peer without a intermediating server is so rare for SaaS that I don't even know of an example.
When inserting text, you can either use mutations that insert the text relative to some reference character(s) or at some index. The former is basically the approach of most text CRDTs. I'm having a hard time imagining the latter without modifying the insertion index relative to concurrent insertions or deletions. I believe that Jupiter based OT imposes a global order, but uses it to determine how concurrent mutations are transformed against one another.
But I don't love how he separates the global order approach from operational transform. Those two approaches go very well together in my experience: - Global ordering eliminates much of the complexity of OT - OT enables optimistic, offline edits and fine-grained merging on top of global ordering
Most SaaS editors are not truly peer-to-peer: edits are sent to a central server before being broadcast back to clients. The service can define the global ordering then - the point of this talk.
But you need the server to be able to reconcile and merge concurrent edits from clients that haven't fully sync'ed yet, and you want clients to be able to locally apply edits before receiving the latest state from the server, so you need some way to undo and replay edits on top of newly discovered edits that happened before (like a Git rebase). This is the basic operational transform function.
The knock on OT has been complexity - the claim that you need n^2 merge functions for n operations. But... most operations in a real app don't collide in ways that require merges so in reality you need far fewer than than n^2. You generally need text and collection merges.
Marijn Haverbeke, the author of ProseMirror and CodeMirror, has a great blog post on a lot of this here: https://marijnhaverbeke.nl/blog/collaborative-editing-cm.htm...
I did, however, create a new collab plugin based on commits from this Google Wave white paper: https://svn.apache.org/repos/asf/incubator/wave/whitepapers/... , and wrote about it here https://stepwisehq.com/blog/2023-07-25-prosemirror-collab-pe... . Anyone needing more performance out of ProseMirror collab while sticking with traditional steps/transactions may find that information useful.
If you're in the area for one of these, I hope you join us! We've got an interesting planned lineup (not all public) among startups, large tech, and academia in the broad systems space.
The same night, Stefan Karpinski of Julia gave a talk on Floating Point Ranges in Julia [1].
For documents, we had to do something more sophisticated. At the time we built our collaborative text editing, the prevailing opinions were that CRDTs were inferior for text editing so we ended up going with Operational Transforms. That took a LOT of work to get right and is still problematic from time to time. If we had it to do all over again now, we'd definitely challenge that decision (and may do so in the future anyway even though we have a system that works).
Thank you for the legendary work on Redis by the way! It's been an invaluable part of multiple large production systems I've built and is such a well-built tool. I'm hoping from your interest in this topic (and some of the blog posts you've written about it) that some part of the roadmap for Redis might help power LLM workloads in the future.
I wonder when VS Code Copilot will start suggesting git rebase solutions.