Cron-based backup for SQLite
litestream.io
litestream.io
Second, Litestream doesn't interfere with other backup methods so you can run it alongside a cron-based backup. I typically run both because I'm overly paranoid and because it's cheap.
Friends don’t let friends use gzip :)
TL;DR, lz4 is 10x faster than gzip (though does not produce as small as a file), zstd is twice as fast and better.
That's it. I'm giving fly.io a try next hobby app I start.
I'm so impressed with their business decisions right now, I was afraid for litestream but then I read how the creator is just hired to work on it full time. What a splendid development!
I actually see this quite a bit. I think the reason for it is that it's a good way to show complexity or moving parts that can break in a process, where the marketing service is just a drop in and start. We can see the SQLite backup process here's and it's not too bad, but we can also see there are a few things that can go wrong that I'm sure litestream takes care of and would allow us to avoid any issues.
For example, backing up the DB onto the local disk before copying it could fail due to lack of disk space, and them we have to deal with the notification and fix for that. I'm sure litestream is a 10 minute setup that handles a lot of intricacies like that.
There is something very powerful about knowing your strength, leaning into it, and building on it.
(Ben might feel competitive about Litestream though; it's his baby. So maybe give him some credit.)
https://www.joelonsoftware.com/2002/06/12/strategy-letter-v/
Reading through the docs I was delighted to find this full set of instructions including script commands for setting up a simpler regular snapshot backup using cron - for situations where Litestream is (quote) "overkill" - ie where the database is smaller and your durability requirements are lower.
In fact, the docs have a whole alternatives section for different requirements. I think this kind of thing is great, and something more projects should do! Wanted to share it because of this meta thing, but also because it's just pretty useful per se.
The best marketing for devs is no marketing. Just being honest, clear and helpful is what works best.
Contrast with so many open source projects that accumulate features and promise the world, so long as you give them a star.
What a cynical take.
Learned that the hard way when implementing sqlite3 backups on Gladys Assistant ( open-source home automation platform https://github.com/GladysAssistant/Gladys )
As I understand it, while you do the backup, other writes should go to the WAL log and only get commited until after the backup?
In my experience, as soon as there is some new data coming in the DB, the .backup command will continue, and if the writes are not stopping, the backup will never stop as well :D
In Gladys case, we put in the application logic a blocking transaction to lock writes during the backup. I haven't found any other way to avoid infinite backups in case of write-heavy databases
> The VACUUM command with an INTO clause is an alternative to the backup API for generating backup copies of a live database....The VACUUM INTO command is transactional in the sense that the generated output database is a consistent snapshot of the original database.
EDIT: Litestream docs will also recommend that: https://github.com/benbjohnson/litestream.io/issues/56
1. call backup_init, backup_step with a step size of -1, then backup_finish. This will lock the db the whole time the backup is taking place and backup the entire db.
2. call backup_init, backup_step in a loop until it returns SQLITE_DONE with a positive step size indicating how many pages to copy, then backup_finish.
With method 2, no db lock is held between backup_step calls. If a write occurs between backup_step calls, the backup API automagically detects this and restarts. I don't know if it looks at the commit count and restarts the backup from the beginning or is smart enough to know the first changed page and restarts from there. Because the lock is released, a continuous stream of writes could prevent the backup from completing.
I looked in the sqlite3 shell command source, and it uses method 2. So if using the .backup command with continuous concurrent writes, you have to take a read lock on the db before .backup to ensure it finishes. It would be nice if the .backup command took a -step option. That would enable the -1 step size feature of method 1. The sqlite3 shell uses a step size of 100.
Another option would be to check backup_remaining() and backup_pagecount() after each step, and if the backup isn't making progress, increase the step size. Once the step size is equal to backup_pagecount() it will succeed, though it may have to lock out concurrent writes for a long time on a large db. There's really no other choice unless you get into managing db logs.
Ben I suggest updating that cron backups documentation page to recommend VACUUM INTO instead!
You can finish the backup in one step, but a read-lock would be held during the entire duration, preventing writes. If you do the backup several pages at a time, then
> If another thread or process writes to the source database while this function is sleeping, then SQLite detects this and usually restarts the backup process when sqlite3_backup_step() is next called. ...
> Whether or not the backup process is restarted as a result of writes to the source database mid-backup, the user can be sure that when the backup operation is completed the backup database contains a consistent and up-to-date snapshot of the original. However: ...
> If the backup process is restarted frequently enough it may never run to completion and the backupDb() function may never return.
The CLI .backup command does non-blocking backup IIRC so is subject to restarts.
https://deadmanssnitch.com/plans vs. https://healthchecks.io/pricing/
The article mentions calling a “dead man” service and that is fine but Cronitor (built by 3 people, no VC dollars) is a proper monitoring solution for jobs not just a “dead man” alert. Happy monitoring
This intuitively feels right to me, but I'd be interested to understand the mechanics behind it. What can go wrong if you ignore this advice and use "mv" to atomically switch out the SQLite file from underneath your application?
My hunch is that this relates to journal and WAL files - I imagine bad things can happen if those no longer match the main database file.
But how about if your database file is opened in read-only or immutable mode and doesn't have an accompanying journal/WAL?
With SQLite and WAL, you need to replace both the main file and the WAL (/delete the WAL at the time you move the backup in place). A simple `mv` won't do that.
It's rather primitive, but could be considered an intermediate step between this blog's full copying and litestream's WAL-based replication.
One on GitHub which is just a snapshot, a single commit in an otherwise empty repo, force pushed. This is for recovery purposes, I don't need the history and would probably run afoul of their service limits if I did so.
And the other on Azure DevOps which has the entire commit history for the past few years. This one is a bit trickier because the pack files end up exhausting disk space if not cleaned up, and garbage collection interrupts the backups. So it clones just the latest commit (grafted), pushes the next commit, and wipes the local repo. No idea how this looks on the remote backend but it's still working without any size complaints, and it's good to know there's an entire snapshotted history there if needed. As well as being able to clone the most recent for recovery if GitHub fails.
Am I deluded in this?
If you're using the default rollback journal mode, it works by copying old pages to a "-journal" file and then updating the main database file with new pages. Your transaction finally commits when you delete the journal file. Copying the main database file during a write transaction can give you either a subset of transaction pages or half-written pages.
If you're using the WAL journaling mode, it works by writing new pages to the "-wal" file and then periodically copying those pages back to the main database file in a process called "checkpointing". If you only "cp" the main database file then you'll be missing all transactions that are in the WAL. There's no time bound on the WAL so you could lose a lot of transactions. Also, you can still corrupt your database backup if you "cp" the main database file during the checkpoint.
You could "cp" the main database file and the WAL file and probably not corrupt it but there's still a race condition where you could copy the main file during checkpointing and then not reach the WAL file before it's deleted by SQLite.
tl;dr is to just use the backup API and not worry about it. :)
But a traditional `cp` will go from left to right in the file, it won't take the contents all at once. Which is explicitly documented as a thing that will break Sqlite - https://www.sqlite.org/howtocorrupt.html
aws s3 cp /path/to/backup.gz s3://mybucket/backup-`date +%H`.gz
The % is going to be interpreted as a newline, it should need escaping as in date +\%HBut it’s in backticks (subshell) so now I’m doubting myself and thinking maybe the subshell saves this but my brain is screaming NO NO BUG! :-)
I’m going to have to spin up a VM to test it and seek inner peace here :-)
> Do not use `cp` to back up SQLite databases. It is not transactionally safe.
I’m curious what happens. Does this mean that it may cause side effects from transactions that are not committed to become visible?
I guess if you could snapshot the database file and WAL and etc. simultaneously this isn’t an issue, right? Because otherwise it would be a problem for a database any time your program or machine crashed.
You have a 70GB database file with "apple" near the start and "banana" near the end. You start cp and it copies apple. The application updates apple to "cat", and then banana to "dog". Finally cp copies dog to backup.
You now have a database which contains "apple" and "dog" which never existed at any point in the original timeline.
I’ve never done this with a SQLite database, but I have done it with running VM file systems (with a CoW backend). And the filesystem pause was always very short.
Do not use `cp` to back up SQLite databases.
It is not transactionally safe.
What if you disable the journal like this:PRAGMA journal_mode=OFF;
Can cp be used to backup the DB then?
And do you still have to send "COMMIT" after each query or will each query be executed immediately then?
If you cp the file, you might end up with different chunks from different points in the transaction history that don't make sense combined together. If you use ".backup", it guarantees that the data in the copy will be consistent with the data as it existed at one of the fsync() calls.
Turning off the journalling will likely increase the chance that your copy will be inconsistent, as there will be more churn in the main data file.
I've read that this is true, but it has always confused me because I would expect that using cp would be equivalent to if the application had crashed / you lost power.
If your filesystem supports snapshotting, it would be safe to create a snapshot and `cp` the database off of that snapshot.
I'm not sure whether this is faster or safer than `.backup` which produces an actual database file with indexes and whatnot, but a benefit of the plain text output is that it's very flexible. The output can be piped to gzip or any other program on the fly, even across the network, without waiting for an intermediate file to be written in full. It can also be imported to other types of RDBMS with minimal changes.
The benefit to using the backup API is that the file is ready-to-go as the database and can be started up immediately.
sqlite> PRAGMA user_version=42;
sqlite> .dump
PRAGMA foreign_keys=OFF;
BEGIN TRANSACTION;
COMMIT;
sqlite>Anyone know if comitting the sqlite db file to git is not safe either?
If you cp a database mid transaction you will loose the transaction as it is processed in the separate journal or WAL file until commit. If you copy mid commit then you will have incomplete data in the main database file and possibly data corruption.
Curious, why is this?
Edit: filed an issue on the docs repo: https://github.com/benbjohnson/litestream.io/issues/54
Pun absolutely intended.