Downsides of Offline First
rxdb.info
rxdb.info
The promise of CRDTs is that unlike most conflict resolution systems, you can layer over a crdt library and basically ignore all the implantation details. Applications should (can) mostly ignore how a CRDT works when building up the stack.
The biggest roadblock to their use is that they’re poorly understood. Well, that and implementation maturity. Automerge-rs merged a PR the other day which brought a 5 minute benchmark run down to 2 seconds. But by bit we’re getting there.
I agree with your assessments here - CRDT is the way forward for most applications; no user wants to fiddle with a merge UI or picking versions like with iCloud. I think RxDB’s position here is from their CouchDB lineage.
> The biggest roadblock to their use is that they’re poorly understood. Well, that and implementation maturity.
I certainly have more understanding to do. My biggest open question is how to design my centralized server side storage system for CRDT data. To service writes from old clients I need to retain a total history of a document, but I don’t want to force new clients to download the entire history nor do I want these big histories in my cache, so I end up wanting a hot/cold system; and building that kind of thing and dealing with the edge cases seems like more than 100 lines of code.
It seems like the Yjs authors also recognize that CRDT storage on the server is an area to address, there was some work on a custom database in 2018, although my thinking is more about how to retrofit text CRDTs into my existing very conservative production cloud software stack than about writing to block storage.
For prose text, what do you think about combining a document-scale CRDT, with fine-grained locking — e.g. splitting the document into a "list of lines/sentences", where lines have identity, and then only allowing one person to be modifying a given line at a time?
I've always felt like this was under-explored, given that in prose text it's almost always semantically incoherent for multiple people to be trying to change a single sentence in different ways at the same time anyway (i.e. they would each have a series of edits they want to do to line A to turn it into either A' or A"; but any result that's not either purely A' or A" will very likely be nonsensical. One could say that the A → A' transformation a user does to a sentence is intentionally transactional.)
I almost thought Notion would be a good example of this, but apparently not (https://www.notion.so/Real-time-collaboration-20a1873baf334d...) — they actually do allow multiple users to be editing the same leaf-node content block at the same time, and so have taken on the full scope of the CRDT problem.
> but I don’t want to force new clients to download the entire history nor do I want these big histories in my cache, so I end up wanting a hot/cold system; and building that kind of thing and dealing with the edge cases seems like more than 100 lines of code.
Yes, but these needs are able to be cleanly abstracted away on the backend — there are internally-complex infrastructure components (like CouchDB, or Kafka) that expose a pure CQRS model to clients, but internally are doing fancier stuff involving reducing changes onto snapshots and then exposing new CQRS streams that begin history with atomic "introduce all this stuff from a snapshot"-typed 'changes'.
There's also some convergent evolution happening here with non-replay sync strategies in blockchains (which can be more interesting to look at if you care about the serverless p2p Operational Transformation type use-case of CRDTs.)
Also, it leads to a poor user experience.
Also, I’m speaking about locking that’s fine-grained temporally as well as spacially: a line would only need to be locked with a ~10s TTL when a user begins to type in that line. Think of it like the user composing a transaction of modifications to the same line, and then committing it. A lot like typing a message into a chat program. Just the user having their cursor on the line, wouldn’t imply that the line is locked; it would only lock when they start typing.
This is already how group-chat apps work, mind you; if you’re an admin who can edit other people’s message lines, you nevertheless can’t edit someone else’s message line while they’re editing it. But they’re only considered to be editing it while they’re actively typing, and for a few seconds after that. If they go idle, someone else edits their message-line, and then you come back and try to submit your edit, it will be rejected. (Of course, that behaviour makes perfect sense for group chat software, where the only other people who can edit your text are moderators, and so moderation actions should “trump” user actions. In a p2p collaboration context, IMHO adding resistance/intentionality to per-line forking, but nevertheless allowing it, makes the most sense.)
Your suggestion also assumes that the network is reliable. What happens if a user takes a lock and there is a partition? If the document is P2P, there is no central authority; when should the other participants override the lock? How much overhead does that add to the protocol?
The main point is that there is no notion of a central clock in a distributed system; hence “lock temporally” is not precise. Relative to which participant? And what happens when messages are dropped? (Even the lock message might be dropped!) A distributed lock implementation is non-trivial.
https://lamport.azurewebsites.net/pubs/time-clocks.pdf
https://static.googleusercontent.com/media/research.google.c...
https://en.m.wikipedia.org/wiki/Fallacies_of_distributed_com...
If you're using a system where you're guaranteed to have knowledge of what other people are editing at all times, there's really no need to use CRDTs in the first place.
If you want collaboration between people, you have to structure it in a way that makes it a conversation, I believe.
I could almost see an idea that you could pattern it after musicians playing together, but that is a very particular kind of rehearsal that has not been done in any other practice, as far as I am aware. Improv may come close, but even that has very specific techniques that really don't make sense in a CRDT landscape.
> My biggest open question is how to design my centralized server side storage system for CRDT data. To service writes from old clients I need to retain a total history of a document, but I don’t want to force new clients to download the entire history nor do I want these big histories in my cache, so I end up wanting a hot/cold system; and building that kind of thing and dealing with the edge cases seems like more than 100 lines of code.
Yeah definitely more than 100 lines of code. I'm sad to report that in diamond types (my own CRDT) I've spent ~12000 lines of code in an attempt to solve some of these problems. I could probably get that down under 3000 loc in a rewrite if I'm happy to throw away some of my optimizations. Doing so would dramatically lower the size of the compiled wasm bundle too - though the wasm bundle is still comfortably under 100kb over the wire, so maybe its fine?
Regarding history, I have a lot of thoughts. The first is that with the right approach, historical data compresses really well. Martin Kleppman's automerge-perf data set has 260k edits ending in 100kb of text. The saving system I'm working on can store the entire editing history (enough to merge changes from any version) in this example with just 23kb of overhead on disk. I think that resulting data set might only need to be accessed in the case of concurrent changes, and then only back as far as the common ancestor. But I haven't implemented that optimization yet.
And yeah; I've been thinking a lot about what a CRDT-native database could look like too. There's way too many interesting and useful problems here to explore.
The Slack link from https://developers.notion.com works for me, maybe Slack has a DNS issue according to this article? https://www.theverge.com/2021/9/30/22702876/slack-is-down-ou...
The Slack link I mentioned is https://join.slack.com/t/notiondevs/shared_invite/zt-lkrnk74...
The "invite has expired"
It's linked in the footer of https://developers.notion.com/
The one issue with CRDT that I find is rarely mentioned and often ignored is the case where you've deployed these data structures that include merge logic to a set of participating nodes that you can't necessarily update at will. Think phones that people don't update, or IOT/sensor devices like electric meters or other devices "in the wild".
When you include merge logic – really any code or rules that dictate what happens when the the data of 2 or more CRDTs are merged – and you have bugs in this code running on devices you can never update, this can be a huge mess. Sure you can implement simple counters easily (like the ones I linked to), and you can even use model checking to validate them. But what about complex tree logic like for edits made to a document? Conflict resolution logic? Distributed file system operations? These are already very complex and hard to get right without multiple versions involved and unfixable bugs causing mayhem.
Having to deal with these bugs in the context of a fleet of participants on a wide range of versions of the code, the combinatorial explosion of the number of possible interactions and effects of these differing versions and bugs taken together can really become impossible to manage.
I'd be interested to hear from folks who have experience with these kinds of issues and how they have dealt with them, especially if they are still convinced that CRDTs were the right choice.
> When you include merge logic – really any code or rules that dictate what happens when the the data of 2 or more CRDTs are merged – and you have bugs in this code running on devices you can never update, this can be a huge mess.
This is a really important point, and its a problem with distributed systems in general. Moxie Marlinspike (the guy behind Signal) gave a great talk on this a few years ago thats well worth a watch. He argues that distributed systems will always be outcompeted by their centralised counterpart because they can't add features as quickly:
https://www.youtube.com/watch?v=Nj3YFprqAr8
We've seen this playing out in real time watching git slowly upgrade its hashing function. The migration is taking years.
I don't know a general solution, but a partial solution in my mind is that we need to knuckle up and make these base layers be correct. Model testing and fuzz testing can absolutely iron out bugs in this stuff, no matter how complex it seems. Bugs in application software usually aren't a big deal, but CRDTs fall in the same category as databases and compilers - we need to prioritise correctness so you can deploy this stuff and forget about it.
Its not as bad as it sounds. The models which underpin even complex list CRDTs are still are orders of magnitude simpler than what compiler authors deal with daily. Correctness in a CRDT is also very easy to test for because there just aren't that many edge cases to find. You can count on one hand the number of ways you can modify a list and the invariants are straightforward to validate.
We used fuzz testing for json-ot a few years ago and never ran into a single bug in production which was in the domain of things the fuzzer tested for. (Though the fuzzer was missing some obscure tests around updating cursor positions.)
If truly peer to peer, that is a lot less clear - do you end up in a collaborative p2p document model forking the documents between ‘new rev’ clients and ‘old rev’ clients? Who ‘wins’? What is the consequence of losing?
At least an API server can clearly reject the client and give an error message - if it’s a CRDT, how does that work?
Just the first example that pops into my head: edit A sets an invoice status to paid, edit b changes the invoice amount from 100 to 120. The merge is a paid invoice with an incorrect amount.
A workaround would be to record a separate PaidInvoice that wont be changed by the application logic.
But that's just a really trivial example that only involves scalar fields inside a single object, and also relies on the application logic considering all ways that the CRDT might behave.
There are countless ways to end up with data that violates the constraints of the domain.
Is there any theoretical groundwork happening on how CRDTs can preserve domain semantics?
> edit A sets an invoice status to paid, edit b changes the invoice amount from 100 to 120. The merge is a paid invoice with an incorrect amount.
There is no such thing as "paid status" of an invoice. That's a simplification. If you unroll it, what you get is accounts that must balance. The invoice starts with "accounts payable"[0] set to 100$, and no payment covering it. Edit A adds a transaction from company account to AP to the value of $100. Edit B changes the liability in AP to $120. The result of a merge should be a transaction record for $100 and liability of $120, meaning a partially-paid invoice.
--
[0] - Or whatever it is called, I keep getting confused by this terminology, as it's not something regular people use in my part of the world.
I don't know anyone who's building a CRDT like that; but its definitely possible. Basho's RIAK databases had a CRDT implementation that would provide a grounding to handle cases like that. When concurrent changes happen, the database simply stored each conflicting value. It was up to the application to decide how the conflicting values merged together.
On top of something like that, you can imagine a bunch of different merging strategies which would be appropriate in different situations: Optimistic (recursively merge all children), Arbitrary-writer-wins, or Conflict-and-bother-the-user. In this case, generating a conflict state is appropriate. Conflicts also make sense when merging branches in git; and I'd love to see a CRDT which managed to reproduce that functionality.
But no, I don't know of any work happening in that space. Maybe the inventor could be you!
Yes, this is the key! 1. store everything, 2. give the domain modelers an array of strategies to merge particular cases, with the default being conflict-and-bother-the-user. Every existing merge strategy can be represented on top of this model, so it should be the base layer. The conflict resolution strategies can be sliced up, layered, hidden, or exposed to the user according to the design criteria because conflict resolution is part the problem domain.
I would go so far as to say that assuming one particular merge strategy for all data structures as 'correct' is actively harmful. What is the worst case outcome in the example above? That nobody looks at it. In a general case like this choosing any default action at all is the wrong call because both edits were made with different assumptions and contexts, and that's the level where the conflict should be resolved, not down here in the invoice.
If you want a decentralized CRDT, you'd probably end up re-inventing smart contracts (which is what I found myself doing once).
Effectively using CRDTs requires a fundamentally different model of state. Typical CRUD stuff simply doesn't work. If you're unwilling or unable to refactor your state layer, then CRDTs are gonna seem totally infeasible.
My current understanding is that CRDTs merely guarantee derministic merging of updates to some basic data structures, where deterministic means that the outcome is well defined and always the same regardless of the order in which updates are merged.
That doesn't mean the outcome makes any sense at all in terms of the sort of application level requirements and constraints you would typically find in a transactional database application. Conflicts may still arise on that level.
So I think what we need to really fulfill the promise of CRDTs is a way to express those application level constraints on top of them.
And his demo implementation (and annotated fork):
* https://github.com/jlongster/crdt-example-app * htps://github.com/clintharris/crdt-example-app_annotated/blob/master/NOTES.md
I wonder why there isn't some open source engine based on this at least for CRUD apps since it has a lot of potential and it is really "simple" to implement and even understand.
After a quick reading the "CRDTs go brr" and the wikipedia page, I think CRDT gives us a mathematical strategy for resolving conflicts. It doesn't mean that the end result will make sense.
The Wikipedia article gives an example of merging an event flag represented by a boolean variable. So the var in this case means that "someone observed this event happening". So the rule for merging this var from different sources is simple, if any source of data reports the var as true, the merged result should be true as well.
The implication is it matters what the data represents, not just whether it is a boolean or a string, etc.
I'm guessing that a colloborative notes field, or a "did someone call this customer" boolean might benefit from a CRDT more so than keeping track of bank account values.
There's Yjs (with a Rust port that is in progress), Automerge-rs, and your own diamond-types project :).
Is Yjs still the current go-to for most projects' needs?
Diamond types will beat yjs on a lot of benchmarks when I have it all working, but I'm still missing important functionality (like loading and saving, and data types other than plain text). The automerge devs have been working on incorporating a bunch of the changes I talked about in my blog post, but they have more work to do before those optimizations have all landed. If I were starting a new project today, I'd use Yjs.
But I suspect in a year or so from now there'll be several high quality CRDTs available to choose from. The space is heating up.
Please do share your experiences that support your opinion.
Gun.js can seem ‘magical’ but then again the 100 lines of code (iirc it’s less) are there for all to see.
I’ve found the time invested to learn about gun (the sync protocol) and RAD (storage abstraction layer) quite worth it. The addition of SEA (security encryption authorization) wrapping around gun enables a lot of implementation options which some find frustrating as they’d rather not have to think about topology before building ‘an app’. Could that be a reason for your comment?
Would love to hear why you think it’s snake oil.
Trac (project management) does some stuff behind the scenes stored in svn, which it uses for edit history. I always liked that idea. Why have two? I have wished for some time that someone did a Trac for git. I just don't want to be the one to write it.
I've also wished for some time that someone would make a new git, not designed by barbarian cannibals, potentially based on CRDTs.
Somehow these notions melted together and now making progress hinges on whether there's a CRDT out there that's up to the task of managing source code - and written in a language I can handle. So far those two haven't appeared, and based on the number of corner cases I've heard described when I watch lectures on CRDTs, I'm pretty sure if I tried to write one myself I'd never see daylight again, but might see the inside of a padded room. Do you think it's safe to say we're still in the distillation phase of invention where CRDT's are concerned? Is this accidental complexity we are seeing or is most of it intrinsic?
I'm instead spending a lot of my free time on the hobby instead of on writing collaborative tools and/or accidentally writing a Trac replacement.
What is the largest-scale or highest-profile real-world usage of CRDT today?
(I glanced at this CRDT vs OT topic before but I'm not up-to-date on where things stand in the real-world performance: https://news.ycombinator.com/item?id=23988999)
It isn't enabled by default but it is very easy to use and the CRDT backend is basically hidden away.
More info:
- https://hexdocs.pm/phoenix/presence.html - https://dockyard.com/blog/2016/03/25/what-makes-phoenix-pres...
It's crossed my (short attention span) radar a couple times and seems interesting. I really love the idea about being able to use it, but also "forget" about it while developing.
Does "leaky abstraction" apply here? If you constrain things enough, can a dev really use it and forget about the details?
https://braid.org/meeting-8 (the shelf part with Greg starts at about 43 minutes in to the recording.)
The code is all here. Its tiny:
Don't make the same mistake as the Google Wave team. Collaborative editing is a great intellectual challenge to work on, but in reality users don't care much about documents that auto-update in real time. In fact, it can be annoying.
Native apps dont deal with the same issues around storage, and are actually _much_ more performant overall. If you haven't done an offline-first app, I highly recommend it. The experience is magical. You can fly around your app at the speed of the users touch. Content is magically loaded as soon as its clicked on. Its amazing as a user.
As for similarities, conflict resolution is a universal problem and is either the most difficult or second most difficult problem [1] to solve. What makes this more difficult is that there is no one-size-fits-all for this. You need to have a deep, nuanced understanding of your system and what makes sense for your use case in terms of resolution strategies. Then you need to implement them, which is not easy especially if your backend is bog-standard REST on a classic SQL datastore.
I've enjoyed reading the past 2 rxdb articles on this (the one mentioned here as well as the one from a couple days ago [2]). Its great to have more content on this publicly available, when I was getting into offline-first I only had a couple options.
[1] - At the end of the day, offline-first is a whole bunch of caching, and so you end up needing to deep dive into cache invalidation strategies which we all know is a hairy problem.
This is a synthetic problem on iOS, created by apple to push users away from webapps. It’s not only safari, apple has completely blocked any browser engines other than their own.
Yes, a lot of commenters justify this by saying they want the web to be documents, and everything else native apps, that every page doesn’t have to be an SPA.
For me it’s the opposite - I really don’t want to install another invasive app, just gimme a webapp. But sadly, apple crippled both storage and notifications enough to force developers AND users to use their appstore.
This is lock-in, against the interest of their users, for their marketshare. The exact same tactics that MS was crucified for.
Stuff like https://github.com/aerogear/offix seem to be in the right direction of what I'm looking for but not nearly mature enough.
I don't want to pu to much effort on the app so I would like something more or less ready made, preferably with graphql apis.
Any suggestions welcome.
The "better" option is to have an offline-aware client-server component (e.g. Realm) which is the primary store as well. This eliminates the sync and so all conflict resolution stays in the same system with well defined semantics.
I wish Couchbase were more helpful practically than they try to present themselves theoretically. Even if their products weren't so expensive the impedance mismatch between their version of CouchDB's sync APIs and their own APIs seems to increase by the year, and is pretty noticeable in how different it works from a PouchDB standpoint and how easy it is to break sync. (Impedance mismatches in allowable database names and _id keys are huge on their own that have massive repercussions in application design.)
Even CouchDB is not CouchDB anymore with impedance mismatches of its own between versions. On the one hand it's good that Cloudant upstreamed a lot of their cluster management tools directly into Apache CouchDB 2+ (even as they made their PaaS offering IBM Cloud only [or whatever it's name of the month is]), but huge architectural changes below the covers in CouchDB 3+ start to present their own sync issues akin to but distinct from Couchbase's (and even some of Cloudant's as they seem [?] to be diverging again into their own 2+ fork after all that work upstreaming stuff?).
More than ever, Azure CosmosDB's focus on bare minimum MongoDB capability and not supporting anything like CouchDB sync, despite having close the same raw ingredients (Cosmos' change feed looks a lot more like Couch's, but is just missing a couple subtle things to make it directly and immediately useful for Couch replication) seems like a "CouchDB is dead and not worth supporting" signal from Microsoft.
Unfortunately, I think PouchDB <=> CouchDB replication has past the "mature" point to the "decrepit" and "falling apart" stage, maybe to the point of "evolutionary dead end" if I'm feeling strongly pessimistic enough, and I've been for years trying to figure what to replace it with.
It's nice to see "we're trying for internals that make it easier to run it at scale", but it's not really a "at scale PaaS" play if there aren't any vendors on the major clouds that aren't IBM's interested in actually running whatever those internals are as a service.
I'm not really into the managed db service offerings, I like to have control over it that they rarely offer.
My tentative solution for the application is to go with PouchDB on node.js on the backend too.
Just giving options and google fodder.
Your users will be frustrated.
Its funny you have to say this - there's a whole generation of programmers who just think web first.
Looks like you're not truly getting the concept of Offline-first apps then. You don't have or need a cache. You have a local database on the device that syncs up with the server, if online.
Conflicts can occur, but modeling data well with a CouchDB/PouchDB setup (i.e. prevent conflicts by not modifying docs all the time, rahter create new ones) + having a simple time-stamp based heuristic can already be sufficient. Otherwise, using CRDT is an option.
The source of truth is the server. Everything else (i.e. the local copy) is a snapshot of that, aka a cache. Its just that offline-first is _always_ a cache-first read, so you seem to think that this makes it not a cache any more, but a regular data store.
The tricky part is conflict resolution.
Yes. In an true offline-first approach with CouchDB/PouchDB, the client-side database can be considered to be main data store and the Server-Side DB could just be a backup. Or it might not be needed at all/ might only be used to migrate from one device to another.
I'd say whether it's a master-slave or Multi-master model depends on the conflict resolution strategy.
- The 7 days IDB limitation does not apply for apps that are published through the stores
- Conflicts can happen, but depending on your design they might not matter in practise. „Implement a proper conflict resulution strategy“ has been on my „todo: maybe“ list for over 3 years now but was never important enough.
- Data migration is not needed as long as schema changes are additive (new doc fields, new doc types). Design carefully early on, keep track of „abandoned“ properties and you'll rarely need a difficult migration.
- Depending on the performance of your customers' phones and the amount of data your app is processing, it (JS -> ... -> IDB and back) might not be fast enough. I had to add caching layers for some use cases. But at some point, you probably want a proper state management library anyways which should include caching nearly for free.
- You can (and should!) still consider most of your data relational. There is even a relational-pouch plugin. But I'm strongly missing foreign key constraints and better DB-level data validation than CouchDB's design docs provide.
I would also like to see a synced version since most browsers these days support syncing settings and passwords. But creating a generic syncing solution that is actually useful is hard.
I think that's "request permission to use files on disk". This is in progress as https://developer.mozilla.org/en-US/docs/Web/API/File_System..., though there's still more work before there's a version all the browsers like (Mozilla likes the ability to work with files, but thinks cross-site access should not be included: https://mozilla.github.io/standards-positions/#native-file-s...)
This is a really difficult problem to solve, and I get Mozilla's hesitation. I'm also frankly very hesitant about Google leading the charge on this, not because I'm paranoid about them sneaking in tracking, just because I think Google tends to create less thoughtful web specifications sometimes.
But... cross-site file access is really important for data portability and open standards, and Google's current proposal isn't bad, it might be rough but it's definitely workable. Mozilla really should try to figure out a way to move forward on this.
We've seen the difference in data portability between mobile and desktop apps, and a big part of the difference between those two platforms is being able to very easily have multiple sources working on your data at the same time. Siloing data has downsides. It's tough to embrace a Unix-style philosophy without allowing programs to operate on the same data. And having Unix-style smaller webapps that work with each other is a good way of fighting against data silos and in some cases a good way of fighting against anti-user and anti-privacy services in general.
I'd love to see more progress made on this, but who knows how that will work out. Caution is probably warranted for the moment, I'm just disappointed that the language suggests Mozilla would never consider a proposal that included this.
It's also very important that this expose user-accessible file system access and not just a virtual filesystem in the browser; otherwise it just becomes another data-silo in the web browser. This is something that Google's proposal really gets right, and it's disappointing to see what appears to be pushback on the idea that users should be able to open up the directories that a web browser is writing to, inspect the files, and open them or move them around the filesystem, or even write to them from native apps. That to me is an essential part of the proposal.
I use CouchDB installed on the client side to implement Offline First data storage. This works for Desktop PCs and if you run it on local device that's accessible via your local network, like a Raspberry Pi in-house for example, you can also use it with mobile devices in-house too.
The local CouchDB will sync with the Cloud based CouchDB as soon as they're both online. CouchDB will decide which version to keep and deliver it from that point.
It's certainly not perfect and doesn't provide "real time collaboration" but so far that's been out of reach anyway and may not be a very good approach at all. The notion of several people editing the same document at the same time seems to me to be chaotic no matter how you approach it.
The biggest downside to this approach is that the user has to install and configure CouchDB. I made a simple web app to help with this, but it's a bit too much to expect users to install and configure it.
What we need is a client side DB pre-installed that any web app can access and the data for that app is sandboxed and can only access the DB assigned to it. But it's not reasonable to add that to a web browser. CouchDB can do that now.
https://developer.mozilla.org/en-US/docs/Web/API/StorageMana...
If you rely on that, and write important things on the browser storage, a week later your user access some moronic site that doesn't work, and their support tells him to clear his browser cache, your data will almost certainly be cleaned with it.
The problem is that browsers do not even consider that they may be storing important data. There is no clear way to recover or backup that data, and it is tangled with what comes from every other site.
I've built a number of apps in the last year or two that use browser local storage. It also annoys me that the idea around "storing data in your browser" is automatically attributed to ad tracking. I have to go out of my way in my apps to inform users that yes this app uses localStorage/cookies, but no this is not used for any ad tracking, rather for actually storing your app's data.
* it can be deleted at anytime (by browser, or even by user!)
* you generally want the server to be authoritative. if there's a bug client-side, server view of state should win.
* it's not possible in the general case to store all user data offline, it's always a subset.
Once you realize that the client-side state is a cache, potential uses of it become a lot more clear.
Sure, you probably want to put some sort of syncing on top, but that isn't even always necessary.
"offline-first" (terrible name, but here we are) generally refers to a classic web application that wants to be able to run offline either for network resiliency reasons or for performance.
"local-first" is a term that has been coined for something close to what you are talking about: https://www.inkandswitch.com/local-first.html
As the user generates new data, spray it to your servers as well as writing it to the syncable IndexedDB local storage, and to an in-memory buffer. Make your backend handle writes idempotently, and retry all failures a few times. (Eg, IndexedDB disk might be full or flaking out, so retry writing the memory buffer to disk.)
As long as the write path is quick, users can tolerate the browser nuking offline storage cache because they can re-download all the data that made it up to your server.
Hopefully soon the browser vendors will allow more durable file system access with appropriate user controls. Chromium built out the file system access API (https://web.dev/file-system-access/) but it’s not supported in Firefox or Safari.
I know terrible, awful bugs eventually doomed WebSQL from getting any traction and IndexedDB seems to be a more competent replacement, but the fact that Google is leaning on FSA seems like a non-starter.
It just feels like there's no way in hell Webkit will ever implement this stuff - not because of the divide between the App Store and PWA's - but due to the implications for privacy.
Hard pass.
On the subject of WebKit, IndexedDB bugs are also pretty bad especially on iOS; we have debated about turning off IndexedDB write buffering in Safari and just do in-memory there. The best thing to do on Apple platforms is to make the app they’re trying to force you to make. Then you can make a little adapter so your web app can write to disk using SQLite and enjoy a nice relational API without needing to worry about the whims of the browser.
If it's the only real option and it's half baked (performance-wise), fine, but I'd be really concerned if the replacement for buggy code didn't actually do what was promised.
On the plus side, SQLite seems to be pretty stable on iOS, so at least there's a chance of it working out.
Edit: Here's the stack overflow copypasta I used to achieve it: https://github.com/EamonnMR/Flythrough.Space/blob/master/src...
I think this is going to be the next "Realm" that works everywhere.
If you have more than 500 users, the price is $500/mo (and it goes up from there).
I started on a design and literally within a few weeks Firefox announced it was deprecated.
I know that Chrome is pushing for a filesystem API but I don't know if that will be exempt from the usual ephemerality. IIRC it is just a private storage space with a filesystem-like API.
i keep doing 'save as' to create new files because i don't trust it lol
Here is the link : https://fr.slideshare.net/ltearno/easing-offline-web-applica...
I do not miss GWT but I think it was inspirational.
My offline-first apps I settled on using ULIDs, which have time stamps and put them first so that they sort lexicographically (such as when included in string CouchDB/PouchDB _ids), and I've been pretty happy with that. That timestamp up front in the first few bytes can help a lot in debugging/"human understanding" of about where the ULID fits in a log stream.
I can also tell you at this point way more than you care to know about storing ULIDs in Microsoft SQL Server to get decent clustered index behaviors.
I think makes a diference if exist a master database or if is peer-to-peer..
Offline first as a principle is more important than web apps. If today's browsers have trouble with offline first, consider that a downside of today's browsers.
Because of Apple, there has to be a severe handicap somewhere, so that nobody can build real cross-platform apps in the browser. That's why iOS browsers can't use APNS/push notifications or reliable local storage, but iOS apps can. Yes, every other browser supports these things, even macOS web browsers. [1]
[1]: https://developer.apple.com/notifications/safari-push-notifi...
This is especially true, because often customer support for many websites have customers begin troubleshooting with clearing the browser's "cache and cookies". In other words deleting all of the local data for all sites. There are ways to delete the local data for just one site, but they are pretty hidden, and involve multiple steps. I wish browsers had a simple "delete all data for just this site" button.
I'm excited about https://github.com/jlongster/absurd-sql though.
Discussion here: https://news.ycombinator.com/item?id=28156831
I trust Mozilla's surface reasons here: they inherited the mess that was NPAPI from Netscape, then decades of experience with XUL binary components, were among the many dealing with Flash bugs and zero-day fallout well after Flash's "heyday", and have combined multiple decades of experience in what happens if the web depends on specific binaries to do its job. From that standpoint of they were already knee deep in trying to sandbox/reign in NPAPI, remove XUL, and remove Flash I absolutely understand why "you want the web to depend on the bugs and zero days of SQLite directly with no abstraction layer between?" was a complete non-starter.
I mean, read all these comments on both articles. People say offline first does not work, then they say that every software is offline first by default and it is nothing new.
It was totally worth it spending two 3 days on it :)
I feel like you missed the most important point of "offline first" is that your data belongs to you and does not need to be shared to the cloud in order to have tremendous value. The code/logic to enhance your data should be shipped to you, rather than the other way around.
Yes, I totally missed this side of the dice. Having your data stored locally, being in control, being able to reset the state to it's origin, export the local state as json and import it somewhere else, knowing what is stored and how long.
> And as for everything related to software, I am reading again what was good is bad, and what was bad is good
Before baselessly accusing someone of writing and submitting clickbait, maybe spend a little more time reading the content and evaluating the context.
Still pretty cool since I'm not a native developer and RN is something I've dabbled in but don't use daily.
To be clear I don't know if the bug was from IndexedDB it was something about length exceeded.
I did try to use small images too eg. some 150px by 150px but going to base64 usually multiplies the size by 1.3
For anyone interested [1] not a phenomenal app but one I poured idk 2-300+hrs into (was working on an RN version too) and went nowhere sucks. I was contributing to one of those codefor# deals.
[1] https://github.com/codeforkansascity/tagging-tracker-alt-app...
Still it's cool to store everything as text/blobs locally.
Trying to engineer your app to work offline causes complexity in the design and implementation for a issue that might never be an issue for most customers.
Assuming your app is pretty crippled when it can't access the cloud.
What's next making your app still work if there is no display by screen reading?
But what if the app is only about storing and displaying data entered by the user and you'd definitely also want to be able to use it on an airplane, e.g. a todo/notetaking/journaling app? Then offline first can make sense.
I'd even say that the complexity of developing an offline-first app can be lower than that of a classical client/server app + caching logic. Sure, you need to figure out how to do schema migrations and you probably want a kill switch to lock out older apps at some point. But the same applies in a classical setting when API-endpoints should be changed or removed. Basically, your document model is now your API. And a very simplistic conflict resolution strategy like „just take the most recent version“ is often good enough.
Once that + basics like auth and account creation are set up, it's very productive and low overhead to work with PouchDB/CouchDB and offline first:
- no need to coordinate with a backend team because nothing else than auth happens there - simpler state management and error handling because the state in the local DB can always be assumed to be correct and the DB is always there
- no need for schema migrations as long as you only add new docs or extend existing ones
- it's great for quickly hacking a prototype or for beginners who just know some HTML/CSS/JS
Any thoughts on this in practice?
They'll fight it with subtlety until the end, and I hope they burn in hell for it.
Encourage developers to make apps that all look/feel completely different in terms of style/UX - no thanks.
I use some native apps and they are great and all, but to be honest I don't care. I want a good user experience and not technical details what's behind the hood of the UI.
I also feel like the stance is pointless since electron apps are here to stay and it just makes the experience of everything worse.
Honestly VS Code is one of the best electron apps available and it still feels clunky compared to basically any native macOS app.
Solved: File system in the browser plus network distribution - https://github.com/prettydiff/share-file-systems
And I'd add than even if all data is generated and stored on the device, syncing capabilities are desirable if the user wants to use his Offline-first app across devices.
The pure native approach only solves (kind of) the limited storage issues, nothing else.
If you nuke your site and only develop about native apps then you'll lose users to competitors because people can try their service without downloading an app.
Yeah maybe we shouldn't have shoved an app run-time into a document viewer, but here we are and things run decently well (thanks to v8 and other tech) if you do things right.
Syncing can be implemented in billion ways, from dedicated solutions to custom ones.
Paradoxically, I think Microsoft got it (somewhat) right with Office. You can use a native C++ app, and this is what most people use because of speed and control. However, if you like working online, don't mind saving your documents unencrypted on Microsoft's systems, don't worry about being offline, you need collaborate with others online, or just like the convenience of having your documents always online, you can use the 365 version with inferior but probably acceptable performance - in this way everybody is happy.
Back in my day we had a word for software that always worked without an internet connection. We called it "software" and it was installed on the user's computer.
I didn't downvote your comment but in this author's article, it's deliberate for the starting context for discussion to be a networked collaborative app.
Yes, internet connected apps is a subset of all possible software but that's not the point.
As an analogy, imagine if someone else submitted an article about C Language memory techniques on an embedded chip. E.g.: https://www.embedded.com/memory-allocation-in-c/
And then a commenter misunderstands that article complaining, "it's confusing to me because this article maintains that the only option is C in an embedded app but over here, I'm using Python with Cloudflare Workers"
In other words, it doesn't seem like you're interested in collaborative apps that require distributed data consistency so this article looks "wrong" to you.
We have an opportunity right now to do an awful lot better. When we do, I want to service both native and web apps. If we do it right, from the network level the distinction should just boil away anyway.
Essentially I want to build an opensource/hackable notes/tasks/calendar system with a central datastore and message broker for coordination between various systems/scripts/components/clients.
The thing I'm struggling with is to define which features are supported in online and offline mode. Every feature that gets added to offline mode adds tons of duplication and complexity. I'm almost at the point of just saying I don't need offline-mode except for viewing already-cached data and maybe very basic creates/updates, with all the processing happening on the backend once the client comes back online.
edit for more context: The old-style native apps (for example OmniFocus) usually have all the logic only in the client and use "dumb" cloud storage for synchronizing between clients. The difference with what I'm trying to build is that I want that central hub to be "smart" and always online so it becomes easy to hack/interact with the system from cronjobs/scripts/external services.
The data model can be reused across many applications, since there's nothing application specific about it. So we can make standard, interoperable debugging tools, backup tools, viewers, etc.
If you want to interact with the same data from scripts, cron jobs and external services, just have another peer on the network with access to the same set of API methods the application can access. You can already read and write data via that API, and any applications with the data open should see any changes instantly.
Basically, what I'm imagining is pretty similar to a self hosted firebase. Except, ideally, I want a CRDT under the hood so we don't need to send all edits via someone else's computer.
A lot of these problems go away if you don't run in a browser, because user inherently trust software more when they're running a copy that can't change from underneath them on a second-by-second basis. I'm a lot more willing to give filesystem access to an application I have to explicitly install and that remains what I installed until I knowingly and intentionally upgrade it, as opposed to code pulled continuously from the network as I am working.
Offline first is only used in the context of web applications and sometimes their Android/iOS cousins (which probably share the same backend, where both are available), once the decision had been made that a not-locally-installed and/or remote synced application is desirable where possible, so isn't being suggested (directly) as an alternative to locally installed offline programs.
If you're looking at a web "site" on a vertical phone screen, one side or the other did something wrong.
My point is: sites and services should have two separate experiences, depending on screensize. These are: desktop site, mobile webapp.
RxDB software comes from this context - it’s a JavaScript library built for this world that attempts to retain the distribution and sharing advantages of the web, while adding back the responsiveness and availability of traditional installed software.
I think you’ll find the original “local first” manifesto more aligned with both the user & traditional installed software with less of a web focused bias: https://www.inkandswitch.com/local-first.html
> In this article we propose “local-first software”: a set of principles for software that enables both collaboration and ownership for users. Local-first ideals include the ability to work offline and collaborate across multiple devices, while also improving the security, privacy, long-term preservation, and user control of data.
Now your user base is roughly split 40/40 between iOS and Android with a non-ignorable 20 running some windows tablet thing. Then for extra fun your access to those users is mediated by the giant black boxes staffed by assholes that are the iOS and Play stores.
And of course back then nobody had any expectation that your software would easily sync between devices and users because that just wasn't something that was doable easily.
The world changes. Yeah it was better for us devs back then, but honestly the new world has some real advantages.
For me personally, I'm cautiously optimistic that PWAs are our way out of the hell that is supporting 3 native platforms for all but the rarest cases.
my impression these days is that web engineers have outnumbered other software engineers. that's a problem because context matters as you are alluding to.
we shouldn't use terms like "offline first" without an appropriate context. or assume context.
Offline first is not "software" as you classified it.
Your "software" you are on about is more accurately described as "offline only", little to none of the functionality requires network access.
Offline first refers to how online functionality is needed for the design of that app, but with steps taken to ensure that even without connectivity for periods of time, it still functions.
Maybe as a slogan or sales pitch...
Reminds me of Fujitsu's Vision Statement: Human-centric Computing. I asked a dozen employees of that company what it meant, and not one could give a coherent answer
- offline-first
- human curation instead of algorithms (or at least transparent algorithms that are customizable)
- the user is not the account number. the account is just the mech that the user climbs into. user can have multiple mechs.
- leverage existing social fabric to provide better user experience. account recovery, etc.
0. Sync with remote
1. Edit files until I'm ready to check-in
2. Stash changes
3. Sync up with latest from remote
4. Pop changes
5. If there are any conflicts, deal with them here, locally. Possibly delete all changes and redo. Redoing is equivalent to someone checking in something at step 0. If this takes some time, move back to Step 2.
6. Push changes
If done this way, the only benefit of CRDT's is during step 5. One of the lessons I've learned and truly believe is that if something sucks, and it gets worse as the problem gets bigger, you need to do it more often. Git merges are a great example of this concept.
And this is where CRDT's are in trouble. The best way to make step 5 easier is by making step 1 smaller. CRDT's viewpoint is that we are offline, and therefore should allow any number of edits, and when we reconnect the system should be able to work it out. It flies in the face of smaller commits. The more changes you need to merge, the harder it is. We live in a world that's connected 99% of the time, and CRDT's simply aren't needed.
On the other hand, the "Offline First" model is great! A user having access to all of their data is wonderful. As the blog notes, it's not preferable for a user to have the entire Wikipedia or Google indexes on their device. So you need to have a use-case for Offline, and we need to do better about this type of stuff, but we don't need to wait on CRDT research for this. Smart caching and materialized views are where I think the real progress is going to come from, making things like Offline Wikipedia possible.