Litestream – Opensource disaster recovery and continuous replication for SQLite
litestream.io
litestream.io
Unfortunately, our org is being pushed away from SQLite because our domain is highly regulated. We have a really hard time getting something like this through technical review, especially for our up-market clients. The only thing I could defend is if this was somehow part of the first party solution and code base.
I can sell SQLite to a bank if they are tolerant to single node failure semantics (I.e. vm snapshot restore). That part is no problem. Already did it half a dozen times myself. It gets tricky if they want more aggressive RPO. I cannot sell log-replicated SQLite if it relies on another, lesser-known 3rd party for that capability. I can get away with murder on my front end vendors but data vendors never seem to escape scrutiny.
I want to hate SQL Server, but it's starting to grow on me. The knowledge that I won't have to argue with a hundred+ CTO/CIOs is starting to win out over the principled nerd things.
Replicating SQLite feels like a cursed thing for many contexts - If someone isn't happy with periodic snapshots, then the nature of their business is probably pretty serious and they probably want deeper (perhaps legal) guarantees about how things would go.
So, running litestream on the node and pushing to S3 isn't a good story, eh? I would have assumed that everything in a single AWS tenant could be a viable path.
I think it's a fantastic story. But, once you bring people who are responsible for billions of dollars worth of assets and PII to the party, things get a bit tricky.
In these settings, I cannot make the argument against the incumbent hosted RDBMS providers if we need a geo replica for business continuity purposes. Hypothetically I could push for it and maybe force it through, but then I'd be left with not very much political capital to force the next thing. I'd also have to be a little bit dishonest about certain aspects and would on the hook for something that not as many engineers could assist me with.
Reliable persistence of data is but one small problem for us and I won't die on this hill anymore. Simply getting everything into the same tenant is table stakes. The hard parts are reducing the # of moving pieces (complexity) and 3rd party vendor count.
Again, for many domains I think litestream can be an easy win. Gaming, AI/LLM state tracking, lesser-regulated SaaS, more flexible RTO, etc. No reason to pay (with money or time) for something more sophisticated than dumping WAL in S3 if you genuinely don't need it.
> Reliable persistence of data is but one small problem for us and I won't die on this hill anymore. Simply getting everything into the same tenant is table stakes. The hard parts are reducing the # of moving pieces (complexity) and 3rd party vendor count.
Yep. It is sad that the perception will be that MS SQL is more simple or Postgres is more simple (I love postgres, but just adding in a network means a lot of potential issues) than SQLite and WAL.
It feels like the argument is really: familiar complexity vs, unfamiliar complexity. And, unfamiliar complexity will never win out no matter how much better the true complexity story is, especially in enterprise domains.
Would you mind sharing some numbers of what you consider an "aggressive RPO", and what RPO number sqlite+litestream can still handle? Of course, it will depend on the use case, but if you would do a rough estimate. I think it could be helpful for a lot of readers.
There are other bank systems that do utilize hard, synchronous replication semantics (hopefully for obvious reasons), but they can afford to sacrifice performance and availability during business hours (more than we can). We always defer to the records in these systems.
We can tolerate a few bits of work getting out of sync and needing to be recovered in the back office. Mopping up a little bit of extra trash after the apocalypse is acceptable to all involved. We cannot tolerate all users being blocked for more than a few minutes, nor can we tolerate a de-sync of aggregate system state exceeding the same. Anything beyond this and we are at risk of having to wipe all work and disrupt end customer activity to keep ops and support teams from crashing.
I'm in a very similar situation in my industry. Thanks for your words.
The difference is that LiteFS is meant for high availability and low global latency, whilst Litestream is designed for disaster recovery.
I'm all-in on server-side SQLite (2022) - https://news.ycombinator.com/item?id=37613747 - Sept 2023 (161 comments)
Cron-based backup for SQLite - https://news.ycombinator.com/item?id=31386330 - May 2022 (82 comments)
LiteStream support for SQLite live read-only replicas - https://news.ycombinator.com/item?id=31007426 - April 2022 (15 comments)
Why I Built Litestream - https://news.ycombinator.com/item?id=26103776 - Feb 2021 (176 comments)
My questions are around that. Are there good solutions for migrations? Has anyone made it relatively easy to handle in terms of deploy/scale/manage?
I've used it in a few places myself, but I've not yet released a non-beta version of it. I'm close to having the confidence to say other people should use it too!
The pitfalls are:
1. It's SQLite, with everything that entails. The schema migration story is "SSH into your production machine and run your migration scripts."
2. Every deploy to a new machine involves downloading your database from S3. Zero-downtime deploys aren't possible because you need to pause writes, sync, and then download the DB to the new container or whatever. If your database is gigabytes (mine is kilobytes), this could be a deal breaker.
Honestly, if managed databases were as cheap as SQLite hosting, I'd just use a managed database, because the SQLite ops workflows are too weird. But I'm confident in the reliability and performance of the site, and it is genuinely cheaper to not run a managed database, which is important for hobby projects.
(Standard internet comment disclaimer that I am a fallible human who might have missed something or am somehow doing it wrong...)
I have no affiliation with Litestream but I was convinced that SQLite could be a viable db option from this great post about it called Consider SQLite: https://blog.wesleyac.com/posts/consider-sqlite
Using SQLite with Litestream helped me to launch the site quickly without having to pay for or configure/manage a db server, especially when I didn't know if the site would make any money and didn't have any personal experience with running production databases. Litestream streams to blackblaze b2 for literally $0 per month which is great. I already had a backblaze account for personal backups and it was easy to just add b2 storage. I've never had to restore from backup so far.
There's a pleasing operational simplicity in this setup — one $14 DigitalOcean droplet serves my entire app (single-threaded still!) and it's been easy to scale vertically by just upgrading the server to the next tier when I started pushing the limits of a droplet. DigitalOcean's "premium" intel and amd droplets use NVMe drives which seem to be especially good with SQLite.
One downside of using SQLite is that there's just not as much community knowledge about using and tuning it for web applications. For example, I'm using it with SvelteKit and there's not much written online about deploying multi-threaded SvelteKit apps with SQLite. Also, not many example configs to learn from. By far the biggest performance improvement I found was turning on memory mapping for SQLite.
Happy to answer any questions you might have!
My goal is also to try and create an app that I don't expect to immediately be successful so it has to be cheap to run in the long term!
I'm not sure what the actual takeaway is. Software that don't require all the manual configuration steps to work properly is better, but it's not always feasible (in this case it's requirements for how the app uses SQLite, which cannot be controlled by Litestream directly). Or maybe Litestream should fail with hard errors if it detects issues in the config (not sure if it's all possible to detect). Or maybe we need to have higher expectations of developers reading the docs properly?
"Stop building slow, complex, fragile software systems"
....but do use data replication over a network.
"Safely run your application on a single server"
...ignoring the fact that this really has nothing to do with "safety" in terms of either security, redundancy, reliability, or durability.
"No-worry backups - Continuously stream SQLite changes to AWS S3, Azure Blob Storage, Google Cloud Storage, SFTP, or NFS."
....ignoring the fact that that's a copy, not a backup. (If you think those are the same thing, you should read a book on backups)
"Runs as a separate process so you can integrate into existing applications with no code changes."
So now you've got to manage two apps, not one? I thought this was simpler?
"Object storage is cheap so there's no need to waste money on additional servers."
If the whole idea is to remove pieces to make something simpler and more reliable, wouldn't you want to use Serverless? Since that's literally zero servers? (CGI apps are the 2000s version of serverless, but way simpler... yet nobody is talking about CGI either)
You could copy the files onto a RAM disk on another computer and call that a backup, but you'd probably be nervous about the computer shutting off and losing your backup. Copying files to a remote disk may have the exact same probability of failure. You don't know until you investigate and calculate the probability. And if you haven't tested the recovery recently (or ever), you don't know the probability of recovery.
The probability of failure or recovery has to consider all eventualities. For example, it's now common practice for cybercriminals to attempt to erase backups of data in a ransomware attack. But it's also possible for people to accidentally delete backups. A proper backup should defend against these eventualities in order to increase the probability of a successful restore.
Your "replication over the network" remark makes little sense to me - adding this kind of replication doesn't make this system any slower or more fragile. The restore action may be fragile, but the system itself works the same as without it.
> So now you've got to manage two apps, not one? I thought this was simpler?
It does make it simpler - there are zero changes to codebase. The "other app" is an addon/a sidecar, which is an extremely common practice. Compare that to creating your own replicated database in a classical sense. Also not all people deploy their own creations. There are plenty of self-hosters who don't code the system they deploy, or don't have that ability/will to be it's developers as well.
> wouldn't you want to use Serverless? Since that's literally zero servers?
How is serverless zero servers? Serverless is (usually proprietary) solution to shared servers, not zero servers. It's also only simple if you don't care where and how your code runs and is secured. If you do care, serverless isn't simple any more.
You seem to also ignore that selfhosting exists and is big. That low cost hosting exists and is surprisingly popular. That many systems don't need multiple instances spread across multiple datacenters, when their RTO and RPO are really low and actual successful disaster recovery scenario is just a bash script that will run for a minute or two.
While you are right to point that this solution is not exactly backups (IMO it sits between HA and DR solutions while combining flaws and advantages of both), your other criticism feels very personal.
If the purpose is to simplify, then you would want to remove the components that add complexity, but here instead of removing them, we've simply changed them, and removed some of their useful functionality. Complexity remains but in a different shape.
My criticisms are personal, in that it feels like the page is trying to be willfully deceptive, and that's annoying.
That's kind of the whole point.