BorgBackup: Deduplicating Archiver
borgbackup.org
borgbackup.org
As one of our users have said[2], borg is "the holy grail" of backups as it does everything rsync always did, and produces remotely encrypted backups that the provider has zero insight into.
It also does not have the inefficiencies that the older, duplicity software has.
If you are willing to go without (borg specific) technical support and do your retention with borg instead of our zfs snapshots, there is a special, discounted rate available.[4]
[1] rsync.net
[2] https://news.ycombinator.com/item?id=17408624
inefficiencies that the older, duplicity software has.
As much as I love duplicity, it does sort of suck when you get into larger numbers of files and gigabytes of data. It's so darn slow sometimes it's almost unusable. Good to know borg is better, I've been meaning to check that out!Sadly, it seems like borg has no built-in support for any kind of redundancy [1]
(1) Are you doing any hardening of the borg remote side? The server is pretty complex and it has lots of code paths which can be exercised by a potentially untrusted client.
(2) Is there a way to protect compromised clients (like a cryptolocker)? borg offers "append only" mode, but it does not allow pruning (obviously). And trying to run "prune" using separate trusted connection will commit untrusted changes as well.
Desktop backup solutions have typically a set of functionalities to compare against (block-based, safe on untrusted repositories, with realtime backup, compression-efficient, and optionally, with a functional GUI), that show how open source solutions are strangely lacking, each in a way or another.
> produces remotely encrypted backups that the provider has zero insight into.
If you're implying an untrusted repository, this is not entirely correct. From https://borgbackup.readthedocs.io/en/stable/internals/securi...:
> in a multiple-client scenario a repository can trick a client into reusing counter values by ignoring counter reservations and replaying the manifest
The last discussion on HN was actually significantly more critical: https://news.ycombinator.com/item?id=18952839.
Borg also has no efficient compression, as it doesn't use multithreading. There are multiple GitHub issues on this subject; I've just checked, and they're closed.
Depending on your use model, this may or may not affect you. Whether you use multiple threads to compress a single stream or not probably matters a lot less if you have 15-30 backups going on at the same time.
I understand this is kinda hard with "zero knowledge" encryption, but it is possible and not a wild feature to request from the self-named holy grail of backups soft.
https://borgbackup.readthedocs.io/en/stable/usage/prune.html
And being careful is difficult, because in Borg, once you do any other backup command from a full access account (for example to do pruning) - it will automatically, no warning, go through the log and apply it. You should really really read up on that functionality first before relying on it, the way Borg has implemented it is close to anti-feature.
Regarding compromisation of strictly the server itself, I believe there are commands to check the state of the repository? Isn't that enough?
It mounts the machine to be backed up and runs Borg on that.
The backup machine does need locking down in of itself but that’s a lot easier to do than locking down something public facing
The snapshots are not deletable even with your full credentials.
This is one of the issues I aimed to solve with BorgBase.com[1]: every single repo is its own backup user and can't see other repos. This separation allowed to add some Borg-specific features, like append-only mode, monitoring for stale backups or using a specific Borg version.
Subaccounts even get their own .ssh/authorized_keys file.
It's all very unix centered and command-line focused. There's no web interface for features like this.
Regular borg only offers two options: either pruning does not decrease space used at all; or pruning permanently commits changes, negating any security advantages. Both of those seem bad enough to prevent offering "append-only" commerically?
In your link the cheapest option is 76€, with an 82€ setup fee.
EDIT - I see it now, had read 4 TB for the cheapest option, but it is actually 40 TB !!
Too much for my typical home user needs, though :-)
I tried borg for some time for full backups from / and there were 3 things I didn't like about it:
1. Mounting and navigating a snapshot was extremely slow. Like wait a few seconds for any `ls` to finish.
2. It needed the encryption password for many things I think it shouldn't.
3. The resulting backups are very opaque.
I've since switched to making my backups by rsync'ing to btrfs subvolumes. I snapshot the previous backup and rsync the differences to the new one. I get to navigate each snapshot with the same snappiness as any other directory in the filesystem. btrfs can also instantly tell me which files differ between 2 snapshots, no matter how huge they are. I don't remember all the options that borg offers on synching, but I doubt it has the same breadth of options as rsync. So for me, rsync + COW filesystem > borg.
I am a little concerned with my use of btrfs, but I'm planning to do the same thing using ext4's reflinks for some disk and filesystem redundancy.
- compressed
- deduplicated
- encrypted using asymmetric encryption, so that the encrypting machine doesn't need to know a password
- durable (using par2 or directly supporting repository repair from an independent backup)/signed
- composed of standard open source tools so that if the backup software goes away, everything can still be retrieved
- supports differing upstream data hosting providers
- open source
There's lots of things that fill most of those requirements, but none that tick all the boxes.
- rdiff-backup does all of the above except deduplication; it's full + incremental which requires a management strategy
- restic doesn't compress, doesn't use asymmetric encryption and doesn't use par2 or similar tools for durability
- borg doesn't use asymmetric encryption and doesn't use par2/durability stuff; it's also pretty slow
I know people call borg the "holy grail", but I think we're still a wee way off that emerging, personally -- even though I do really like (and use!) borg.
Whatever the details were, I came away with the distinct impression that the borgbackup developers didn't really respect the gravity of archived records and long-term storage. While that's probably not a totally fair conclusion from a one-shot use, in my defense, I'll offer up the latest version of their changelog [0], where:
* the first thing listed is an apparently serious data corruption bug that lived through several stable releases
* the second thing listed is an apparently-very-serious security vulnerability ("a flaw in the cryptographic authentication scheme used in Borg allowed an attacker to spoof the manifest ...")
* the third thing listed is another data corruption bug titled "Pre-1.0.9 data loss"
* the fourth thing listed is another data corruption bug titled "Pre-1.0.4 potential repo corruption"
Note these are all post-1.0 versions. To be frank, I dare not scroll further.
Users are depending on this software to safeguard important files with the assumption that it will be able to reproduce them bit-for-bit-intact some years down the road. Any long-term storage software requires developers with a fanatical devotion to compatibility, longevity, and integrity if it's ever going to be more than a toy.
I hate to pile on open-source devs who're trying their hand at making software that a lot of people appreciate, but overall, it's hard for me to say I regret the choice to pass it up and stick to combinations of ZFS, rsync, and the trusty old tarball.
[0] https://github.com/borgbackup/borg/commit/75dcf9356334188276...
This is a changelog backup software which does data storage and encryption mainly. If it ever has bugs worth talking about they're going to be in... data storage and encryption. If we look at rsync for a change which is much older, simpler and mature: https://download.samba.org/pub/rsync/src/rsync-3.1.1-NEWS
- fixing .. traversal in V3 (how did that survive so long?)
- issues with copying attributes
(There won't be much data corruption since rsync does 1:1 copies)
Time-to-fix could give us a better idea maybe? Either way, software doing X has bugs in X is "normal".
It's even worse because unlike rsync, which synchronizes filetrees from point A to B and can be immediately confirmed to have either worked correctly or not, borg uses a bespoke storage format that's not easily verified or operated upon by standard utilities. You just have to trust it to pull the correct data out when you need it. That's a heavy burden to put on a tool, and the borg devs seem to be struggling under its weight.
Consider the standards expected of filesystem or database maintainers. borg repos are not really that different -- they're a big, opaque chunks of bytes that you expect to be able to produce specific data on demand with perfect reliability.
These types of systems don't get marked stable when they're still in primordial, eat-your-data development phases, and on the small handful of unfortunate occasions when data loss bugs sneak in, they're a) usually limited to some bizarre corner case; and b) taken extremely seriously, almost to the point of solemnity, and often result in major overhauls to a project's validation and QA routine.
If any stable filesystem or database had 4 widely-applicable database corruption bugs within a year, it'd be lights out for that project. All trust lost, reputation irreparably ruined, angry letters to a variety of mailing lists, permanently-increased scrutiny on any new projects or maintainers among the same class of projects, etc.
All I'm saying is that based on my admittedly-murky prior experience, the project's observed track record, and the high standard of care that must be met to qualify a project for storing authoritative copies of data, I'm personally not comfortable trusting borg with anything important any time soon. No one is obliged to share that evaluation, of course.
I think Backblaze B2 is the cheapest, but it's not supported by Borg. So a backup system without compression support, but with B2 support, may ultimately be cheaper than a system with opposite features.
It's also important to consider the use case. If a user's bulk of data is photos/videos/audio files, there's virtually no use of compression - but of course, the proportion will vary per user.
I'd treat it as a tool in a bigger backup system (which is roughly what rdedup does with it).
`gzip --rsyncable` and friends are mostly useful for when you already have compressed files on your filesystem, and want to sync/backup them byte-for-byte. Of course the benefit is lost unless most compressed files in the world are produced with such mode, and sadly most aren't :-(
The author of `zsync` did various experiments confirming gzip --rsyncable then sync is sub-optimal, and implemented a somewhat crazy "look inside" approach that can sync compressed files byte-for-byte AND very efficiently:
> gzip --rsync does fairly well, with both rsync and zsync transferring about 410kB at the optimum point. zsync with the look-inside method does much better than either of these, with as little as 140K transferred. > -- http://zsync.moria.org.uk/paper/ch03s03.html
IIUC though it's more of a "2-files sync" scenario like rsync, not applicable to "chunk everything then dedup" approach of borg and similar tools.
We have also developed a Qt-based desktop client[2] that runs in the system tray and makes it easier to browse archives or do restores.
For headless deployments, I highly recommend Dan's Borgmatic[3]. You could deploy it all together with our Ansible role[4].
1. remote backup through ssh, especially de-duplicating between different clients, especially concurrent client backups.
borg needs a compatible version installed on the other side (and compatibility has been broken between versions in the past); bup uses bare-bones ssh+sftp, so the other side can basically be anything.
borg will have a lot of download to client from server if multiple clients back up to the same repository (essentially every time in most common use cases); bup will have a minimal download.
borg maintains a repo lock, so multiple clients backing up to the same repo will be serialized; bup does not, so it can be concurrent.
2. Storage format
borg's format is it's own format; bup's underlying data format is basically a git repo (which you can treat as such; you may need to manually apply "cat" to rebuild files, bit "git" and "cat" are all you need).
I have heard good things about restic, but did not have a chance to evaluate it myself.
restic fell over hard at around 100 terabytes; the others at around 500.
I'm backing up a little over 3 petabytes. I use Bacula. It's awful, but I haven't found anything else that can deal with that kind of volume.
Did you test 'Burp'[1]?
What ultimately worked? Plain old rsync over ssh, to a zfs pool with snapshots and compression.
As far as I could figure, the only notable downside of this is that the storage device must be trusted, since it has access to all of the data, and that you effectively needed root permissions on the storage you're copying to for filesystem permissions which makes a multi tenant backup server cumbersome (you could chroot or use containers or something but these solutions can become fiddly, e.g. running multiple instances of ssh on nonstandard ports to enable multitenancy).
EDIT: Ah I see rsync 3.0.0 (released in 2008) fixed this issue. Maybe I was on an old version.
But I also do backups to tape (LTO-8, in a Storagetek library), and I do like that Bacula handles that for me.
The dedupe is downright magical. We've been able to remain super frugal on the storage allocation for backups solely because of how wonderful its dedupe is.
Restores are also really nice and easy since they are just a FUSE mount.
If you have a Linux host or hosts, I can wholeheartedly recommend borg. It's elegant, robust, and fast. And open source.
Compared to the Bacula and BackupExec, Borg was lightning fast and its disk usage was very frugal.
I collected a list here: https://github.com/albertz/wiki/blob/master/backup-software....
It took me some time to figure it all out. I wrote a script to not manually repeat configuration steps next time. Although it's quite opinionated it still grew to 600 lines.
Feel free to check it out, maybe it can help someone https://github.com/senotrusov/backup-script
One limitation that isn't mentioned often is that when doing anything with your borg repo, RAM use increases with the number of files you have backed up. In my case I had 15 million files, and mounting the repo took quite a bit of time (minutes) and used 11GB of RAM.
Restic also has the same issue.
https://www.rsync.net/products/borg.html
What is s3 these days ? 2.x cents ? Plus traffic ? We don't charge for traffic/usage/bandiwdth in any way ...
[1] You do get technical support, just not specific support for setting up your borg backups, which can be fairly complicated ...
You're not ever going to get immediate, personal technical support from a UNIX engineer at those services like you do at rsync.net.
How do you all keep track of whether your backups succeeded? I'd like to receive an email if a scheduled backup didn't run. The only thing I have for now is rsync's feature where they warn you if your data hasn't changed by X kb in the last Y hours/days but I find this lacking because multiple machines write to my rsync account. It'll only warn me if none made any change but I'd never know if only a few failed. Same thing if there is nothing new to backup, I get an email from rsync.net but I don't know if every backup job failed or there is just no changes.
So far S3QL is the least worst of all these deduplicating/compressing/encrypting backup solutions I have tried, but I havent tried borgbackup. Despite it's stated focus on cloud object storage, S3QL also works great on NFS and local filesystems as a target as well.. and sshfs..
[0]: https://chocolatey.org/packages/borgbackup#testingResults
Now, I just installed it on my mac and running `borg --help` takes about 5 seconds before outputting the help.
Same for any other `borg ...` command.
I'm not quite sure why, but that's the only command I've noticed to be running slow on my system.
I'm running borg 1.1.10.
62% C 35% Python
I've never seen borg take 5+ seconds for anything except creating backups. The same `borg --help` command finishes in 0.4s on my machine. That's not blazing fast either but I'm okay with it.
Also, BorgBackup is not new; they have been around for a while.
I didn't think you could trademark someone's last name?
https://en.wikipedia.org/wiki/Bj%C3%B6rn_Borg
A search of borg in the patent office shows the name used a lot: http://tmsearch.uspto.gov/bin/showfield?f=toc&state=4806%3Av...
If you look at https://en.wikipedia.org/wiki/Apple_(name) , the surname 'Apple' is used quite a bit.
I hate how easily corruptible trademark/copyright/patent rules are
But more importantly, it is the name for a castle in several languages. So unless Paramamount is planning to sue a lot of very old places in Norway and other countries...
Borg at Google is not a product per se, it's a internal tool/service. Pretty sure you can name internal tools whatever you want.
As an user of such critical piece of software, I would ask the following questions:
* Is the software popular?
* For how many years it has been battle tested?
* Are there many reports of data corruption?
* Is it well maintained?
* Do the maintainers provide support?
Borg/Attic has been greatly used for many years now and check all the boxes above. I don't think I've ever read about data corruption that was caused by a bug in it.
Which language it is written in is nothing more than a mere curiosity for me.
Emphasis on scratch not on RIR ("rewrite in Rust") ;-)
It would need to respect heterogenous networks (Microsoft Volume Shadow Copy, for example), networks at all, meaning global deduplication, full volume imaging and file-level backups.