Local-first software: You own your data, in spite of the cloud (2019)
inkandswitch.com
inkandswitch.com
I believe local-first peer-to-peer networks have gotten close-enough to being able to deliver feature parity with centralized Application Service Providers for 5+ years now, the industry just hasn't caught up.
For social media platforms like Facebook and X, their "killer app" features are all solvable with hash-based data structures (chains, trees, forests), p2p capability systems, and gossip protocols. Identity systems are solvable in ways that non-technical people can understand.
Part of the solution was the p2p stack maturing. The other part is that centralized solutions have normalized a lot of user flows that p2p can be competitive with, users have been trained up on a lot of patterns that previously wouldn't have been palatable in the market.
There is extreme market demand for it, and the zeitgeist is tuned into the societal and financial failings of ASPs right now.
Facebook is trying to act as the free relay and archivist for the entire world's social graph, there are a lot of hard problems to solve there which personally, I think contributed to social networking silently dying and degenerating into the monstrosity that Social Media is.
Once this stuff starts making it into end-user products, I suspect it's going to be a one way door. P2P can provide user experiences that centralized services can't match - either for technical, financial, or legal reasons.
We are in a sweet spot of opportunity for building the web as it was promised to us.
If folks working in this space want to compare notes, or are looking for work on a well funded no-BS R&D team developing this space, my email is in my bio.
People have been conditioned to believe all software must be free, but that SaaS is fine. People are willing to rent software in the cloud but not pay for it at the endpoint.
Add to this the fact that the cloud is the ultimate DRM. It can’t be pirated if the user doesn’t even have binaries or have possession of their own data.
SaaS provides inescapable recurring revenue and very strong lock in. It’s massively superior from a business model point of view. Meanwhile the (much more) evil twin of SaaS, namely “free” surveillance and addiction driven apps, dominates B2C.
People might want local first software but if they won’t pay for it that doesn’t matter. The whole industry will wrap itself around whatever model pays. Everything is all about cloud and SaaS not because these are technically superior or better for the user but because that's where you get cash flow.
As soon as computing becomes a primary job and not a toy... people clam up about their ideas and stick to safe, generic, bland output.
Computing as-a-job for random plebs, has sacrificed the protection of ideas in pursuit of more innovation and freedom. Resulting in the least innovation and freedom we've seen in decades.
Until we can walk on water again, this spiral will not end. Why people cannot see this spiral and treat it like an existential threat to the web, I have no idea.
A 'new kind of network' is surface level analysis, we need to stop computers sucking the life (and consciousness) out of everything they touch.
Cash flow is way easier, when you can work on something novel and deliver it to the end customer without the concepts being stolen along the way.
Do you have any ideas on how we might do that?
There's a constant tearing conflict in me, between the status quo programming I've learnt over the years and how I'd want it to be.
It's been bubbling up for a while, everytime I touch a keyboard now I'm rapidly frustrated. It's also difficult to know if it could be finished before we switch to TPM/smaller/more locked down computers.
It is hideously conflicting to be working on privacy tech and need to be internet connected to survelliance at the same time, how can you even begin when teacher is peeking over your shoulder constantly?
It feels like death to share it, sorry!
I've got to get it going and then let it live on it's own terms. Anything else is a tainted experiment I guess.
1) The number of people with control over you should be as small as possible.
2) The number of people held responsible for your actions should also be as small as possible.
3) The number of people trusted with your security should also be as small as possible.
If you are going to give up control in favor of convenience you should carefully consider if it is worth it and for how long. Control also means you can pay a huge price at any time. There is a certain risk for the deal to suddenly not be worth it.
If you are going to have others be responsible for your writings and actions you should also carefully consider if it is worth it. The responsible thing to do could be to preemptively shut you down when in doubt. The legal situation/law may also change and require critical analysis of things you wrote long ago. Each word you write is a ticket in the false-positive lottery.
Then when all of the above works out fine you still have to trust people, all of your documents may end up publicly available after some breach out of your control or encrypted behind a ransom paywall. It is going to be different from putting things on usb drives and keeping them in your vault physically.
I am concerned by the threat to open source this model brings to us.
Eg I’m tempted to go with a stack of a centralized Postgres with a SQLite for each user (a la Electricsql, mvsql, etc.) because I want users to have their data local as much as possible (in their own SQLite db) but a centralized Postgres database still feels logistically necessary to coordinate eg user accounts and sharing.
I’m all ears on better p2p solutions. I’ve looked into things like eg using matrix or nostr but they don’t feel like the solution I’m looking for yet. But I’m eager to find something that can keep data local and have sharing done more p2p.
Our current gossip structure for identity is a hash tree. The secure scuttlebutt research paper by Tarr is a good place to start for that.
For capabilities, we are (currently) using the same concepts as UCAN.
I'm sure they're doing a lot of cool stuff - it's just I have no idea what that stuff is.
Y.js and automerge emerged as solutions combining CRDTs and content transfer, they look really promising. There is a Y.rs version if that's better for you.
I've always dreamt of building something on top of Syncthing, ie something that would use file synchronization. It's more versatile and will definitely last longer than anything else, and it has some built-in capabilities for having a third party helping transport but not being allowed to read content.
I recently came across https://github.com/superfly/corrosion , a service discovery and state management tool that is working completely p2p. CR-SQLite, in particular, allows multiple tables from multiple databases to be merged thanks to CRDTs. I'm sure there's a lot to build on top of it.
I feel like you're not really interested in full p2p but want some centralization point to manage some auth stuff, so I'd investigate couchdb/pouchdb first.
For example in my budgeting app:
- All of the user's budget data is stored client side in IndexedDB
- Offline support via a service worker
- Peer to peer synchronization between devices via WebRTC data channel
- Lightweight IdP server which also handles the WebRTC signalling (via web sockets)
This setup is nice because the user's data never touches the server and very little server side infrastructure is needed (none at all if you don't need data synchronization). One downside with the data sync is that both devices must be online with the app open at the same time to sync, but perhaps this could be improved with web workers?This also means that I need to have a constantly running web worker on my phone/tablet etc...? In my time I am also trying to solve this problem of "data never touching a server" but I believe we're not there yet, despite all the devices we have: sometimes the smart watch is off, sometimes the laptop is off, etc., so you need a mechanism similar to "Google drive", and who never had sync issues with that? And that's a supercentralized solution. It's super hard to "merge" in a distributed async system.
[1] https://github.com/actualbudget/actual
[2] https://github.com/actualbudget/actual/blob/d1e57340b88960d0...
Recently there have been CRDT solutions that try to solve this problem, however, such as Triplit [0], or ElectricSQL [1].
For CouchDB an equivalent exists in the form of PouchDB: https://pouchdb.com/
Runs/syncs to the browser too which is just lovely.
See my other reply: https://news.ycombinator.com/item?id=37743517#37746101
It didn't stop the bugs. There are lots of ways to screw up even in a fully online system. It certainly helped though!
I'm working on a few projects in this area:
- https://www.typecell.org - Notion meets Notebook-style live programming for TypeScript / React
- https://www.blocknotejs.org - a rich text editor built on TipTap / Prosemirror that supports Yjs for local-first collaboration
- https://syncedstore.org - a wrapper around Yjs for easier development
In my experience so far, some things get more complicated when building a local-first application, and some things get a lot easier. What gets easier is that once you've modeled and implemented the data-layer (which does require you to rethink / unlearn a few principles), you don't need to worry about data-fetching, errors etc. as much as in a regular "API-based" app.
Another interesting video I recommend on this topic is about Linear's "Sync Engine" which employs some of the local-first techniques as well: https://www.youtube.com/watch?v=Wo2m3jaJixU
I took a look at the landing page out of curiosity, just an FYI but at first glance there's nothing that indicates to me that this is not a regular SaaS app.
Specifically, unless I'm missing something, nothing in the text jumped out at me indicating the app satisfies this condition stated in the article:
>for good offline support it is desirable for the software to run as a locally installed executable on your device
Might want to make this feature more prominent if you support it.
Although it's entirely architected on a local-first stack, I indeed haven't shipped the main benefit of this, a locally installable app. There's a WIP PR here that adds PWA support: https://github.com/TypeCellOS/TypeCell/pull/352. I'll highlight this more when this is merged.
Nevertheless, some of the benefits are already noticeable and come "out of the box" with building on a local first architecture, even if not shipping an executable yet: - multiplayer sync - speed: documents are loaded from local storage initially if they have been loaded before, and changes sync in after that
In the future (when there's an installable app), I also want to enable saving / loading from the file system, so that it's completely transparent where your data is.
https://news.ycombinator.com/item?id=26266881 - Feb 25, 2021 (90 comments)
https://news.ycombinator.com/item?id=21581444 - Nov 20, 2019 (241 comments)
https://news.ycombinator.com/item?id=19804478 - May 3, 2019 (191 comments)
The cloud is a prison. can the local-first software movement set us free? - https://news.ycombinator.com/item?id=36984692 - Aug 2023 (206 comments)
Local-First Software (2019) - https://news.ycombinator.com/item?id=31594613 - June 2022 (29 comments)
Local-First Software:You Own Your Data, in Spite of the Cloud (2019) [pdf] - https://news.ycombinator.com/item?id=26266881 - Feb 2021 (90 comments)
What if we had Local-First Software? - https://news.ycombinator.com/item?id=24790170 - Oct 2020 (144 comments)
Local-first software (2019) - https://news.ycombinator.com/item?id=24027663 - Aug 2020 (131 comments)
Local-First Software (2019) - https://news.ycombinator.com/item?id=23985816 - July 2020 (9 comments)
Local-first software: You own your data, in spite of the cloud - https://news.ycombinator.com/item?id=23966558 - July 2020 (1 comment)
Local-first software: you own your data, in spite of the cloud - https://news.ycombinator.com/item?id=21581444 - Nov 2019 (239 comments)
Local-first software: You own your data, in spite of the cloud - https://news.ycombinator.com/item?id=19804478 - May 2019 (190 comments)
I've built and use a time tracker, a double entry accounting system (using a Beancount(ish) syntax) and an activity tracker. Two of the three of those were using older CLI code I wrote a few years ago adapted into an extension.
Adding a webview that tracks the content of a tab is pretty simple. It also gives me the excuse to write parsers (some of my favorite code) which allows me to render the text into a data structure that my webview code (in my case React) can render on each keystroke. I don't get fancy, I parse the full buffer on each keystroke and with proper React state management the DOM updates are optimized.
I don't think I'll ever open source or publish the extensions themselves as they are too specialized for what I need and want but I may (if I have time) generalize the webview + tab change/rename/delete/close tracking code into an open source library.
I'd recommend this approach if it fits with your side projects. Its an easy way to get a web gui for your ideas without adding to the graveyard of broken domain dreams.
1. CRDTs are a pain, especially rich text. Schema migrations are a pain. And eventually when you have a very successful product, you'll have issues with vector clocks that grow linearly and unbounded with every user who pops by to collaborate. If you're successful enough to have a 2000 person workspace, you won't be able to / want to sync an entire workspace of data in realtime to everyone's local machine, so the whole system devolves into a glorified local cache (nothing wrong with that, but it's not a CRDT and arguably isn't local-first).
2. Freemium SaaS has proven to be the most lucrative business strategy. Users get stuff for free which widens the top of the funnel. Students go into the industry and bring their tools with them, and then you charge enterprises for collaborative workspaces per-seat (I call this the Slack playbook). Given the challenges of CRDTs outlined above, you're making this job really hard.
3. If you're making all of your money on enterprises, they would prefer that a rogue employee can't save all of the company information offline when they leave. That's why you can't export your company Gmail or Notion workspace. Employees will pretty much always be online so local-first isn't really necessary for them, although everyone does want the performance of local-first.
4. There are large swathes of problems that are not suited for local-first -- basically any application that requires transactional writes. You can't enforce simple invariants like "every use must have a unique username". Almost every application has requirements like this unless you specifically design around them. Sometimes, they can't be designed around at all.
At the end of the day, I like to think of choosing local-first through the lens of the "technology budget". If you want to build something local-first, then that's your entire technology budget and then some. Do you want to bet your entire business on that? Or do you want to use tried-and-true tools and spend that technology budget on something else that moves the needle?
Developers love to know how something is built, but users really don't care so long as it works great! Example: Obsidian.md is awesome, and local-first. But it's mostly developers who use it -- everyone else just uses Notion.
I still love Local-First, but as far as I'm concerned, it's a developer honeytrap.
It seems to make a lot of sense for them, which is probably why cloud providers discourage it.
Data is the important part. Cloud compute can be swapped.
The simplest equivalent is just regular apps that save data to somewhere that's backed up.
The masters of this game are Apple. Bar none. All their software is local first and asynchronously syncs to iCloud (optionally!) using CloudKit, which is a relatively consistent replicated database. This design is hard to master but worth starting with. They also support P2P that works (Bonjour and friends). Apple can do this because they're firmly rooted in the software development milieu of the 80s and 90s. They never got the memo about the web taking over and revel in writing high quality desktop apps. Users love them too. Developers like to pretend this ecosystem doesn't really exist outside of mobile, but it very much does.
There are a few things that make this hard for other kinds of developers to do well.
One is that it requires apps to have file formats, and most devs don't learn these days how to design them. They know SQL but not how to wrangle complicated byte arrays. And our infrastructure kinda sucks at working with files. You can't POST a directory over HTTP, for example. SQLite helps a lot here, but you still have to publicly document your schemas if you want the user to truly have a chance at "owning" their data. The cloud then just holds a backup of these files.
The desire for strong collaboration and sharing features is another problem. Apple isn't particularly big on those. However, a lot of software barely needs it! In most cases a way to publish or share documents with a hyperlink is good enough, especially if others can comment on it. A lot of users don't actually want to allow unrestricted real time edits to their work as it undermines ownership, hence why Google Docs has now moved to a more Word-like workflow where edits are help in suspension until explicitly accepted.
The biggest problem is that the only form of sharing that actually works well is an HTTP link, due mostly to OS makers catastrophically dropping the ball on anything else. And that pushes people towards the web. Apple solve this with light web renderers for their file formats that let people view read only copies, more or less. Often this is sufficient.
The Apple approach + an openly installable backend is a reasonable way to do local-first software. It doesn't imply the software has to be free or open source. It can still be proprietary, you just need to sell the backend as software as well as the frontend.
Are there any resources to learn more about the differences it makes for a developer, when implementing a traditional Server-client architecture as opposed to a P2P approach? So far, I have only worked with server-client architectures, it's hard for me to estimate the pros / cons.
Also, is there an option for Automerge to sync via a central server instead? Any examples?
My approach would be to categorize user content into
- CRDT-enabled content considered as the norm, and
- content trapped in legacy containers that do not support CRDT.
With an implied warning for the user: legacy content offers a limited user experience.
Migration from "legacy" to "normal" should be fully supported, and as transparent as possible (but needs dedicated support for each file format). Converting "normal" (CRDT-enabled) content into files/blobs should be treated as a "take an snapshot" export operation.