Duplicity: Encrypted bandwidth-efficient backup
duplicity.us
duplicity.us
The problem I have with duplicity and backups tools of its kind is that you still need to create a full backup again periodically, unless you want to have an ever-growing sequence of increments from the day you started doing backups.
Content-addressed backups avoid that, because all snapshots are complete (even if the backup process itself is incremental), but their content blobs are shared and eventually garbage collected when no references exist to them.
My tool of choice is kopia. Also borgbackup does similar things (though borgbackup is still unable to back up to the same repo from multiple hosts at the same time, though I haven't checked this for a while). Both do encryption, but its symmetric, so the client will have keys to opening the backups as well. If you require asymmetric encryption then these tools are not for you—though I guess this is not a technical requirement for this approach, so maybe one day a content-addressed backup tool with asymmetric encryption will appear?
And if so, what would be the main differences between just committing to a git repo for example?
Better yet would be to use a rolling hash to decide where to cut the blocks, and then use a locality-aware hash (SimHash, etc.) to find similar blocks. Perform a topological sort to decide which blocks to store as diffs of others.
Microsoft had some enterprise product that performed distribution somewhat like this, but also recursively using similarity hashes to see if the diffs were similar to existing files on the far machine.
It's available as a built-in component of Windows, it's just a library with an API.
Essentially the MS RDC protocol is just rsync run twice in a row, with the rsync metadata copied via rsync to compress it further.
There's an important difference is that RDC uses a locality-sensive hash algorithm (MinHash, IIRC) to find files likely to have matching sections, whereas rsync only considers the version of the same file sitting on the far host. rsync encodes differences on each file in isolation, whereas RDC looks at the entire corpus of files on the volume.
For example, if you do the Windows equivalent of cat local/b.txt >> local/a.txt, rsync is going to miss the opportunity to encode local/a.txt -> remote/a.txt using matching runs from a common local/b.txt and remote/b.txt. However, RDC has the opportunity to notice that the local/a.txt -> remote/a.txt diff is very similar to remote/b.txt and further delta-encode the diff as a diff against remote/b.txt.
Here's a demo.
First, create two test files. The files both contain the same two 1-megabyte chunks of random bytes, but in the opposite order:
$ openssl rand 1000000 > a
$ openssl rand 1000000 > b
$ cat a b > ab
$ cat b a > ba
$ du -sh ab ba
2.0M ab
2.0M ba
Commit them to Git and see that it requires 4 MB to store these two 2 MB files: $ git init
Initialized empty Git repository in /tmp/x/.git/
$ git add ab ba
$ git commit -m 'add two files'
[master (root-commit) 7f75af0] add two files
2 files changed, 0 insertions(+), 0 deletions(-)
create mode 100644 ab
create mode 100644 ba
$ du -sh .git
4.0M .git
Run garbage collection, which creates a packfile. Note "delta 1" and note that disk usage dropped to a bit over the size of one of the files. $ git gc
Enumerating objects: 4, done.
Counting objects: 100% (4/4), done.
Delta compression using up to 6 threads
Compressing objects: 100% (4/4), done.
Writing objects: 100% (4/4), done.
Total 4 (delta 1), reused 0 (delta 0), pack-reused 0
$ du -sh .git
2.1M .git
I'm not sure as if it's as sophisticated as some backup tools, though.I think it is a valid way to consider them. Another option is to think of the backup as a special kind of file system snapshot that manifests itself as real files as opposed to data on a block device.
> And if so, what would be the main differences between just committing to a git repo for example?
The main difference is that good backup tools allow you to delete backups and free up the space whereas git is not really designed for this.
Wouldn't that mean that, when using encrypted backups, secrets would have to be shared across multiple clients?
If I'm understanding it correctly, it sounds like an anti-feature. Do other backup tools do that?
I'm not sure if content-addressed storage is feasible to implement otherwise. Maybe use the hash of the the unencrypted or shared-key-encrypted as the key, and then encrypt the per-block keys with keys of the clients who have the contents would do it. In any case, I'm not aware of such backup tools (I imagine most just don't encrypt anything).
For example, Pika Backup, and Vorta are popular UIs for Borg of which no equivalent exists for Restic, while Borgmatic seems to be a de-facto standard for profile configuration.
For my own purposes, I've been using a script I found on Github[0] for a while, but it only really supports Backblaze B2 AFAIK.[1] I've been meaning to try autorestic[2] and resticprofile[3] as they are potentially more flexible than the script I'm currently using but the fact that there are so many competing tools - many of which are no longer maintained - makes it difficult to choose a specific one.
Prestic[4] looks intriguing for my partner's use, although it seems to have very few users. :\ A fork of Vorta[5] seems to have fizzled out six years ago.
[0] https://github.com/erikw/restic-automatic-backup-scheduler
[1] https://github.com/erikw/restic-automatic-backup-scheduler/i...
[2] https://github.com/cupcakearmy/autorestic
[3] https://github.com/creativeprojects/resticprofile
Have you considered?:
https://github.com/netinvent/npbackup
Or (not FOSS, but restore-compatible):
Relica looks neat, but at that point I'd either suggest she uses one of the Borg tools or write a simple wrapper for her to trigger backups instead.
edit: Still looks a bit hairy for an average user to install currently, and the maintainer writes "I'm not planning on full macos support since I don't own any mac" - https://github.com/netinvent/npbackup/issues/28
scheduling+browsing https://forum.restic.net/t/backrest-a-cross-platform-backup-... (Golang webUI supporting Linux and MacOS).
A number of others are findable through the community section of the forum.
Bit of a self plug, I author Backrest. The most significant challenge historically has been that restic has poor programatic interfaces but in recent revisions the JSON API (over stdout) has largely stabilized for most commands.
This is great, when you do your first restic backup on a machine it uploads all the data which takes a long time and if there is the tiniest interruption (like computer going to sleep) then you have to start from zero again, at least that's the experience I had. Instead I went via excluding the biggest directories and then removing them from the exclusion list one by one, doing backup runs in between.
I ask because my Borg repo is an order of magnitude smaller because of dedup, so it's essential for me.
I use it weekly for system as well media backup (yt-dlp for some YouTube content as a hedge in case the channel is ever unavailable in the future).
Basically still an issue. The machine takes an exclusive lock and it also adds override since each machine has to update it's local data cache (or whatever it's called) because they're constantly getting out of sync when another machine backs up
bupstash looks promising as a close-to-but-more-performant borg alternative but it's still basically alpha quality
It's unfortunate peer-to-peer Crashplan died
It has less features than Kopia, but what's there looks like high-quality to me.
(I'm also using it to back up 150 TB (300 million files), on which all other dedup programs run out of memory.)
The documentation refers to files and directories. Does the software let you take a consistent, point-in-time snapshot of a whole drive (or even multiple volumes), e.g. using something like VSS? Or if you want that have you got to use other software (like Macrium Reflect) to produce a file?
Where does the client cache your encryption password/key? (or do you have to enter it each session)
If you have problem using it, please let me know.
I haven't. I use local Ceph S3 for backups, and then use kopia to mirror that to one a local RAID just in case my Ceph dies ;-).
It stores the password, base64-encoded, to ~/.config/kopia/repository.config.kopia-password. I suppose it would be nice, at least for workstations, if it supported keyrings—and it might, I haven't looked into it.
Doesn't this increase the chance of data loss? If a blob gets corrupted, then all the backups referencing that blob will have the same corrupted file(s). This is similar to having a corrupted index in an incremental backup chain (or maybe in this case you would lose everything?), but in the case of incremental backups the risk is mitigated by periodically performing full backups. Also my gut feeling is that you will save space with content-addressed backups only if you're backing up multiple machines that share files, but in the tipical average user scenario where one is backing up a single PC you get a similar space usage. Keep in mind that you tipically delete bacups older than a certain threshold. Could you maybe comment on my points?
It's much more efficient to deduplicate, then add redundancy. Like say storing said blobs on a RAIDz3. Or use backblaze's approach and split the blob into 17 pieces, add 3 pieces of redundancy, and distribute the chunks across 20 racks.
If you are serious of course you'd have an onsite backup, deduplicated, with added redundancy AND the same offsite.
that said will give Borg a look
Borg is now my holy backup grail. Wish I could backup incrementally to AWS glacier storage but that just me sounding like an ungrateful begger. I'm incredibly grateful and happy with Borg!
It's quick, tiny and easy... and restores are the easiest, just mount the backup, browse the snapshot, and copy files where needed.
AFAIK, the only difference is that Restic doesn't require Restic installed on the remote server, so you can efficiently backup to things like S3 or FTP. Other than that, both are fantastic.
Not practical for huge backups but it works for me as I'm backing up my machines configuration and code directories only. ~60MB, and that includes a lot of code and some data (SQL, JSON et. al.)
Do you have a direct link I can look at?
Its front page hints at this, but there must be details somewhere.
I modified the scripts to do `sleep 1` between each change but it left a sour taste and I never gave Restic a fair chance. I see a good amount of praise in this thread, I'll definitely revisit it when I get a little free time and energy.
Because yeah, it's not expected you'll make a second backup snapshots <1s after the first one. :D
So I'll keep an eye on Rustic instead (it is much faster on some hot paths + allows you to specify base path of the backup; long story but I need that feature a lot because I also do copies of my stuff to network disks and when you backup from there you want to rewrite the path inside the backup snapshot).
Rustic compresses equivalently to Borg which is not a surprise because both use zstd on the same compression level.
Borg does have the option to run both a client-side and a server-side process if you’re backing up to a remote server over SSH, but it’s entirely optional.
Duplicati restores can take what seems like the heat death of the universe to restore a repo as little as 500Gb. I've lost a laptop worth of files to it. You can find tonnes of posts on the Duplicati forums which retell the same story [0].
I've moved to Borg and backing up to a Hetzner Storage Box. I've restored many times with no issue.
Remember folks, test your backups.
[0] https://forum.duplicati.com/t/several-days-and-no-restore-fe...
Since you mention it, I am seizing the opportunity to ask: how should borg backup be tested ? Can it be automated ?
borg check --verify-data REPOSITORY_OR_ARCHIVE
You can add that to a cron job.Alternatively, I think the Vorta GUI also has a way to easily schedule it[1].
I'll add that one thing I like to do once in a blue-moon is to spin-up a VM and try to recover a few random files. While the check command checks that the data is there and theoretically recoverable, nothing really beats proving to yourself that you can, in a clean environment, recover your files.
[0] https://borgbackup.readthedocs.io/en/stable/usage/check.html
> borg check --verify-data REPOSITORY_OR_ARCHIVE
Thanks ! I thought there were some more convoluted process but I couldn't picture out anything except extracting the whole archives and check up by hand.
Then just listing the files in the archive is a not-bad way to find an obvious problem. Or straight up unpacking it.
But if you're asking about a separate parity file that can be used to check and correct errors -- I haven't done that.
It also has extensive support for ignoring stuff and it works very well.
I still use Borg because its policy of expiring older snapshots is more useful for me, but Kopia is extremely solid and I would use it any day if I didn't care that it doesn't actually keep one monthly backup for the last 3 months as Borg does (it decides which older snapshots to keep with another algorithm; it's documented on their website).
It's fantastic to have so many great open source backup solutions. I investigated many and settled on restic. It still brings me joy to actually use it, it's so simple and hassle free.
My only complaint is that, like a lot of software written in Python, it has no regard for traditional UNIX behavior (keep quiet unless you have something meaningful to say), so I have to live with cron reporting stuff like:
"/usr/lib/python2.7/dist-packages/paramiko/rsakey.py:99: DeprecationWarning: signer and verifier have been deprecated. Please use sign and verify instead. algorithm=hashes.SHA1()"
along with stuff I actually do (or might) care about.
Oh well.
You are using an old version of Duplicity. It dropped all support for Python 2 in 2022: https://git.launchpad.net/duplicity/commit/setup.py?id=5505f...
# $1 # local folder
# $2 # bucket
declare -a exclude=(
"node_modules"
"Applications"
"Public"
)args=""
for item in "${exclude[@]}";
do
args+=" --exclude '*/$item/*' --exclude '$item/*'";
donecmd="aws s3 sync '$1' 's3://$2$1' --include '*' $args"
eval "$cmd"
I agree this is not the absolute most optimized solution but it does work quite well for me and is easily extendible with other scripts and S3 CLI commands. Theoretically if Borgbackup or Duplicity are backing up to S3 they're using all the same commands as the S3 CLI/SDK.
Besides, shell scripting is fun!
They are not. Both Borg and Duplicity pack files into compressed, encrypted archives before uploading them to S3; "s3 sync" literally just uploads each file as an object with no additional processing.
Differential backup here means that if a file has changed you only send the change delta, not the whole file. This is what makes it possible tobrun that kind of things every hour if needed even on large folders.
If s3 supported that and with the already existing versioning you'd have a pretty kickass solution; that's basically what you can do with rsync and a zfs filesystem for example
That's the most obvious glaring problem, beyond that it's just kind of garbage in terms of the amount of space and time it requires to perform restores. Especially restores of files having many reverse-differential increments leading back to the desired restore point. It can require ~2X a given file's size in spare space to assemble the desired version, while it iteratively reconstructs all the intermediate versions in arriving at the desired version. Unless someone improved this since I last had to deal with it, which is possible, it's been years.
Source: Ages ago I worked for a startup[1] that shipped a backup appliance originally implemented by contractors using rdiff-backup behind the scenes. Writing a replacement that didn't suck but was compatible with rdiff-backup's repos while adding newfangled stuff like transactional backups with no need for "regress", direct read-only FUSE access of restore points without needing space, and synthetic virtual-NTFS style access for booting VMs off restore points consumed several years of my life...
There are far better options in 2024.
[0] https://github.com/rdiff-backup/rdiff-backup/blob/master/src...