Automerge: A library of data structures for building collaborative applications
automerge.org
automerge.org
It was by far the most developer-friendly experience I've had trying to implement collaborative editing. The one thing it didn't have that YJS did was built-in undo/redo.
okay looks like partykit hides some extra metered costs with a call us button
reflect's pricing seems a lot easier to understand!
just trying to think how this will work with cloudflare + fastify
Biggest problem with yJS for me has been the ergonomics when I use it with React. There’s a third party project called synced store that I used for the stream overlay but it has some strange behavior.
With first party support with React in automerge I think it’s worth a shot for my rewrite
It’s great at first, but woefully underdocumented if you want to use it in a way that doesn’t have off the shelf support, and the code is tough to parse (for me at least). Same with subdocuments.
I wanted to use it with Lexical a while back, and yjs plugin was too tightly coupled too a single data model and was too complicated to DIY.
The best way to manage is starting with a raw empty doc, defind all the keys, then update that with a custom blob.
Once you have this doc on your clients, its just a matter of applying the latest updates.
But yeah, Yjs has an eccentric code structure which is impossible to parse without understanding the ops rrquired.
Id look at indexeddb and webrtc plugins
- Yjs is mostly made by a single author (Kevin Jahns). It does not store the full document history, but it does support arbitrarily many checkpoints which you can rewind a document to. Yjs is written in JavaScript. There’s a rust rewrite (Yrs) but it’s significantly slower than the JavaScript version for some reason. (5-10x slower last I checked).
- Automerge was started by Martin Kleppmann, Cambridge professor and author of Designing Data Intensive Applications. They have some funding now and as I understand it there are people working on it full time. To me it feels a bit more researchy - for example the team has been working on Byzantine fault tolerance features, rich text and other interesting but novel stuff. These days it’s written in rust, with wasm builds for the web. Automerge stores the entire history of a document, so unlike Yjs, deleted items are stored forever - with the costs and benefits that brings. Automerge is also significantly slower and less memory efficient than Yjs for large text documents. (It takes ~10 seconds & 200mb of ram to load a 100 page document in my tests.) I’m assured the team is working on optimisations; which is good because I would very much like to see more attention in that area.
They’re both good projects, but honestly both could use a lot of love. I’d love to have a “SQLite of local first software”. I think we’re close, but not quite there yet.
(There are some much faster test based CRDTs around if that’s your jam. Aside from my own work, Cola is also a very impressive and clean - and orders of magnitude faster than Yjs and automerge.)
We have recently published a new research paper on replicating SQLite [1] in a local-first manner. We think it goes a step closer to that goal.
> Convergent, Replicated SQLite. Multi-writer and CRDT support for SQLite
From "SQLedge: Replicate Postgres to SQLite on the Edge" (2023) https://news.ycombinator.com/item?id=37063238#37067980 :
>> In technical terms: cr-sqlite adds multi-master replication and partition tolerance to SQLite via conflict free replicated data types (CRDTs) and/or causally ordered event logs
Braid aims to make it easy for such systems, as they’re built, to be able to talk to each other.
Bruinen.co
Shoot me a note if you want an early build! Or if interested in building with us :)
tevon [at] bruinen.co
Making the app work without an internet connection is step one. Making it reparable without an internet connection is step two.
Step two is blocked if you can't keep the code for all of the app's dependencies near enough at hand such that its still accessible after the network partitions.
This was a thing around 2 years ago. Nowadays speeds is the same or in favor of Rust, depending on the benchmark in question.
For example, in one of my tests I'm seeing these times:
Yjs: 74ms
Yrs: 9.5ms
That's exceptionally fast.
This speedup seems to be consistent throughout my testing data. For comparison, automerge takes 1100ms to load the same editing history from disk, using its own file format. I'd really love to see automerge be competitive here.
(The test is loading / merging a saved text editing trace from a yjs file, recorded from a single user typing about 100 pages of text).
One trick we’ve been pulling is tailing the automerge contents into a sqlite db in-browser for more complex querying.
(some notes on how/why here: https://tender.run/blog/tender-and-crdts)
For non-text collaboration, there is a more crowded "market", because it is an easier problem to solve - at least when your app has a central server. Tools range from hosted platforms like Firebase RTDB to DIY solutions like Figma's (https://www.figma.com/blog/how-figmas-multiplayer-technology...). Meanwhile, Automerge's target niche is decentralized collaborative apps, which are rarer.
A boilerplate including user authentication & authorization
Tech: Automerge, tRPC, Prisma and deployment on fly.io and vercel.com Bonus: includes explanation videos on the website
I love the idea of "local first software" https://www.inkandswitch.com/local-first/
A number of past discussions too: https://hn.algolia.com/?q=local+first
If you have a web client that only does on-client data storage, you're not dependent on centralized server-side storage for persistence (which is the main privacy-loss hazard).
The problem is that the app host still has all the technological freedom to not honor their privacy agreement, and there don't seem to be backstops to that behavior in browsers.
Uses OPFS + WASM to run SQLite and host files. I've had a pretty healthy response to it from HN so I plan to make a desktop app + add PeerJS for file sharing.
It would be great if programs were collaborative out of the box.
I see their being overhead and what not that would make it very domain specific and less appealing for anything that doesn't need collaboration.
People said the same about web programming and yet we have e.g. Svelte.
Some things are just not so well expressed in a framework. Especially things that manipulate state in special ways.
Also, CRDT's don't provide synchronization for free. They ensure that all concurrent modifications will be merged somehow. If the data being synchronized has any structure, it requires careful CRDT-aware data model design to ensure the merging is semantically reasonable (or that, in the worst case, incompatible changes produce a detectably broken state).
There’s definitely some room for interesting work here and language level support could be cool.
Elixir interpreter clustering and otp is maybe the closest existing thing, which is awesome but only for existing erlang fans.
A library or even built-in language support for distributed data structures will take a decade or two to get to the point of proving a set of features to be truly helpful and good/necessary to have as a library or maybe even as a built-in feature.
BTW, we don't even have quickcheck yet, since the generation and proof reduction quality isn't great across implementations :/
collaboration software benefits a lot from this approach because people can just use it and know it won’t get out of sync
Also, anything where you might need data offline at a specific moment. Like if the internet goes down but you still need the thing to show to the guy at the place, you’ll be glad if they used crdts because you’ll automatically have a copy
Also, cheapskates who don’t want to pay cloud bills
but will you ever be able to trust the automatic merge?
Imagine a complex document with a few collaborators which work in parallel offline and make significant changes (that is the use case of automerge?).
You will have to proofread the entire document after every non trivial automatic merge, or how good does this work in practice? Would it be easier to just wait for a wifi connection and do the changes in real time?
You will have to store it in the cloud eventually, otherwise it is just local software, not local-first
With IPV6 and Torr onion routing NAT traversals is almost becoming a non issue.
> While Automerge optimizes for working offline and merging changes periodically, Pigeon is optimized for online real-time collaboration.
Vector clock. Why not a Hybrid Logical Clock?