Update: yet another take on basically the same approach, also as a self-contained binary: https://github.com/kopia/kopia
Update: yet another take on basically the same approach, also as a self-contained binary: https://github.com/kopia/kopia
It contains everything needed (python, python libs, other libs) except glibc and related libs which must match the OS / kernel and therefore are intentionally not included.
restic fails on repo mounted via CIFS/samba on Linux using go 1.14 build https://github.com/restic/restic/issues/2659
Of course there are some cifs issues in the borg tracker, too, but most of them seem to be based on unreliable network connections - while the restic problem seems to be related to golang details.
With backups I would rather not trust a tool that has basic problems like accessing files on a network drive, but this is just me.
Kopia is too new for that state.
Given that all of the options being discussed here look technically better and in some cases more than an order of magnitude cheaper than other popular backup services and software discussed on HN in the past you could afford to run full redundant backups with multiple combinations of software and backing storage and still have more options for a much lower price than a few years ago.
A few highlights (for me at least):
- They have implemented support for object locking and versioning to prevent a compromised set of S3 keys being used to erase backups. There seems to be a pull request outstanding to dynamically shift forward the compliance hold each time an object file is "deemed still used", but very few s3 bucket based backup tools seem to really think about resisting deletion and documenting best practice. I like to use minimum privilege and deny DeleteObject permission, but using versioning and object locking policies seems to be an interesting potential solution.
- Inclusion of (optional) Reed Solomon error correction codes in backup objects.
- a CLI verification option that will test a given % of your backup objects (at random) every time it's run, to give you opportunistic random verification to try find issues.
- reasonably well documented CLI options for recovering data were anything to go wrong - as the developers mention, there are a wide range of ways to recover data. Interestingly, they appear to have actually implemented these, rather than leave them as theoretical. For example, CLI commands to interact with indexes, manifests, blobs, snapshots, etc. You don't need them normally, but having these available lets you look around a simple backup and understand the format and gain confidence in it, and that the tools work.
The only feature it seems to me to be lacking is asymmetric encryption of backups, so you can keep the key needed for recovery away from the host you are backing up.
> - a CLI verification option that will test a given % of your backup objects (at random) every time it's run, to give you opportunistic random verification to try find issues.
FWIW Restic also supports this - see restic check --read-data-subset (https://restic.readthedocs.io/en/stable/045_working_with_rep...)
Also Restic doesn't have any problems with buckets with deletions disallowed as long as you allow them for just the `locks` directory. A policy I used for one of my targets looked like:
{
"Statement": [
{
"Sid": "AllowAdditions",
"Effect": "Allow",
"Action": [
"s3:PutObject",
"s3:GetObject",
"s3:ListBucket",
"s3:GetBucketLocation"
],
"Resource": "arn:aws:s3:::BUCKETNAME"
},
{
"Sid": "AllowDeleteLocks",
"Effect": "Allow",
"Action": "s3:DeleteObject",
"Resource": "arn:aws:s3:::BUCKETNAME/locks/*"
}
]
}
And it even works with buckets with object locking enabled since >= 0.13.Then pruning could be done once every few months. The egress fees could be high though.
https://help.backblaze.com/hc/en-us/articles/217667478-Under...
So I settled on borg. I use it for offsite backup. I did restores of individual files as well of whole snapshots after broken hard drives. There even was a time I used it with WSL to backup my parents data.
I suppose the latency of the remote server was too high.
Except that I tested them and they didn't. I hadn't "lost locally cached metadata". And it didn't just scan all the source files again, it re-transferred them over the network. All 4TB of it! That's why it took so long. (At that time the sever was still local, because I wanted to make the initial run not over the Internet. So there wasn't even added latency.)
I'm sure it was a bug, maybe even in combination with OpenBSD, which moste likely is fixed now. As I said it was years ago. I compared and was more happy with borg. Which stood the test of time and saved my ass multiple times.
`restic backup -v -r <location>/<repo_name> <source_dir>`
first `cd` to the parent directory of the `source_dir` to avoid too many nested directories in the repository and conflicts in the mounting point path (so that the file change detection of restic works when mounting the drive to be backed up at another point)
The performance issues you experienced may have been resolved.
Is that also true for Kopia?
We rely on this in some places in fact. We have a couple of redundant hosts all serving the same file set. This deduplication allows us to push backups from all systems without coordination and only the first backup of the same file set requires storage beyond metadata.
This may make sense, depending on your threat model.
And yeah, we tend to partition our borg repos along functionality and needs of restore. For example, database backups and file store backups are split into different repos, but several file store hosts in the same cluster all write to the same borg repo. After all, if something allowed to compromise one file store host, it will most likely allow compromise of all identically setup file store hosts in that cluster.
And on the other hand, in this way, we don't have to think about the backups if one of the file store hosts goes offline - the other two will just continue writing backups. This would be something that's really easy to forget and could bite very badly.
> Another difference between the two programs is related to deduplication. Borg is designed with the assumption that each machine being backed up will use its own repository. Letting multiple machines backup to the same repository can impact performance, and simultaneous backups from different machines to the same respository are not supported.
So I should have said "Borg has limitations" instead of saying the support is absent for deduplication across machines.
When you say "without coordination", does it mean simultaneous backups from several machines to the same repo are possible?
It is, with a bit of (other) coordination. Basically, only one process can actively write to a borg repository. Once something is writing to the borg repository, the repo is locked and nothing else can write to that repository.
However, you can easily stagger backups from several hosts into the same repository - node 1 writes at 09:00, node 2 writes at 10:00 and node 3 writes at 11:00. This works without problems, it just needs some monitoring for quickly growing backup sets in case your timing goes awry. You can also configure a wait timeout, how long a borg process will wait to lock the repository, but that will require some tinkering with SSH heartbeats to avoid connection timeouts, as borg won't talk over the wire while waiting.
As the documentation correctly states, this slows down the backup writing process a bit, because each node has to synchronize its local chunk cache at the start of a backup or a prune. This wouldn't be necessary if each node had their own dedicated backup repository. This can take 5ish minutes on our large repos (4TB+) and usually takes less than a minute for our smaller repos (<1TB). It's not really a big deal imo.
However - and we made that mistake earlier - something like `borg check` and other commands become really slow if you have some 15TB - 20TB+ borg repository and borg repos that large become really messy to manage. It works, but some things work at a glacial speed on a good day - while locking out all other backups. That's why we tend to group our borg repos by dataset and storage type ("All databases supporting app X backup into the app-database-X repository"). This way, all databases are in that repository even after failovers or switches to geo-standbys and it's easy to setup a restore procedure. Hoqever the overall size of the repository is somewhat limited and manageable.
Borg can can run in WSL but has seen limited testing under such, per their own docs.
That's bad if you want to use it
1) on NixOS (I don't want backup configs laying around in `~/.config`). As Indy famously said: "That belongs in a Nix expression!"
Edit: On a second note: Why have mandatory config files at all? I don't have an issue with having the option or it being the default, but for my use case being able to specify the whole repository config via arguments sounds considerably more sane.
2) with ZFS snapshots (yes, I'm backing up `/path/to/dataset/.zfs/snapshot/<timestamp>/foo/bar`, but that should not be its path in the metadata!)
OTOH, it seems to have the upside that you can apparently alter snapshots after the fact more easily (e.g. if you find out you shouldn't have backed up that gigantic VM image you just moved somewhere temporarily). I leave the decision on whether this is a footgun or not to you.
And to be clear: The ZFS snaphot thing is also a pain with Restic, too. You can hack around it somewhat better with something like systemd-nspawn, but it really shouldn't be that hard.
Backup tool authors never seem to support this use case.
Instead of working out how to teach my backup tools about snapshots, I just mount them in a subtree and use that as a chroot env.
Looks like compression only added in the latest release of restic.
No mounts on windows.
No GUI?
Lack of compression was likely the reason I went with Kopia.