Restic – Backups Done Right
restic.net
restic.net
And some resources on how they're different:
- https://github.com/restic/restic/issues/1875
- https://stickleback.dk/borg-or-restic/
- https://sysadministrivia.com/episodes/S4E5
The general concensus seems to be that restic is borg with more whistles (backing up to various places), but borg is the more trusted tool with the longer history (just use SSH and be done with it). I personally recently used borg for a migration between computers and it worked great for me.
If anyone though knows how I can more easily restore a file that would be great. I have to supply the restoration snapshot ID right now, and I'd rather just do 'latest' and have it find the newest version of the file in all of the snapshots and restore it. Is this possible ?
Something like:
restic restore -i '/path/to/file.txt' -i . latest
instead of: restic restore -i '/path/to/file.txt' -i . search-file-snapshot-list-for-big-sha-here restic mount backup/
and pick the file(s) you are looking for from the `backup/snapshots/latest/` directory.A couple weeks ago I finished a client-side encryption module, got another weekend's worth of work to integrate encryption support into the back end. The client-side encryption module is already in the Git repository, but there is no documentation for it yet. My next project is a Web based GUI (maybe an Electron based client-side GUI too, depending on what makes sense).
Oh, and the most recent release supports granular user permissions, so you can grant a host (via an SSH account on the backup server) permissions to create backups, but not delete them. Or have an administrative user that can expire old backups but not read backups, for example. Can help with thwarting crypto viruses that try to delete backups.
So am I right to infer that there is a ridiculous amount of overhead in the data transfer, as opposed to some of the other software mentioned? Would you use this over the internet? Or an untrusted network connection?
For data encryption (the main part is completed, but need to expand the backend to recognize encrypted data, should be finished shortly) -- the output of "tar" is piped through "tarcrypt". What tarcrypt does is it takes a standard tar file input, compresses/encrypts the file data, and outputs a tar file with some extended headers that contain info about the compression/encryption, including the RSA public key fingerprint used to encrypt the data, an HMAC, and the encrypted (passphrase-protected) private key (this can be made optional). The idea is that the encryption itself is AES-256-GCM, with a random key, which is encrypted with RSA public key. That way you can have encrypted backups without needing to have a password sitting in plain text on the client. And the RSA private key is passphrase encrypted, and sent along with the tar header to the server. On restore, you will be prompted (client-side) for the passphrase. This way you can restore a client even if the keyfile is destroyed.
I plan to make the encrypted key storage optional, but that would require that you manage the key file backup separately, and doesn't get you much more security (assuming you have an adequately strong passphrase).
Server requirements are a server with ssh access, and the snebu binary installed (optionally suid to a non-privileged backup user account, so that granular permissions can be employed for other accounts). And since Snebu is written in C, with only liblzo2, libcrypt, and sqlite2 as dependencies, it is easy to get it to work with a wide variety of systems (and the client side only requires a modern enough version of GNU "find" and "tar", unless encryption is used, which would require "tarcrypt" also -- modern in this case means withing the last 10 years, the "find" command needs to support -printf with the appropriate parameters).
Borg runs doing incremental backups to a local directory. I sync the borg backup folder to a free BackBlaze B2 bucket. The whole thing comes in around 5GB of backup including database backup and all the hosted files and configuration files.
Other differences which some might consider good and others bad include:
* Borg is in Python, Restic is in Go? Python is a more widely known language, the chance that more developers will be able to maintain it is higher, the chance for new features is higher, the chance that bugs will be found is higher. Both still provide a static binary download so deployment is very easy for both.
* Borg allows setting names for archives (which can still be templated with things like {now} or {username}), while Restic automatically names them with random ids. Borg's approach seems more user-friendly.
* Borg allows more choice in what encryption is used, including authenticated/non-authenticated, and including using no encryption at all (might be useful in some cases). (https://borgbackup.readthedocs.io/en/stable/usage/init.html#...)
* Borg allows specifying "max_segment_size = xxxx" on a repository, thus allowing to somewhat adjust the size of archive ("segment") files while they are stored on the server (default being 500 MB). Restic uses its own sizing and (if I remember correctly) more often than not it produces very small files (a couple of MBs) which crowd the filesystem. In general, archive file sizes can affect the filesystem and the uploading them to the cloud, where many cloud providers API operates on individual files, not on parts of them.
flagging any directory for exclusion directory exclude with a tag file inside it: It has this. It's called --exclude-caches and will exclude any directory with a "CACHEDIR.TAG" file in with the correct content.
I can't think of any reason why you would need to customise the name of the file, but I'm sure there are use cases out there.
I might like to mark certain things as not backed up because they contain sensitive or otherwise unrelated information. But I might not want other tools treat them as cache data, that is, transient and safe to delete.
--exclude-if-present "CACHEDIR.TAG:Signature: 8a477f597d28d172789f06886806bc55"
You can specify the tag filename, and optionally append ":initial content" (the tag file must start with this content to avoid false tagfile matches).
[1] https://restic.readthedocs.io/en/stable/040_backup.html#incl...
One doesn't need an interpreter + supporting libraries to run a backup program. With a Go program one just drops a statically compiled binary in $PATH and it just works.
Anyway, I'm still using duplicity. It backups on rsync, ftp, you name it. Written in Python as well, unfortunately. Don't care abot S3 etc as I don't like loss of control over my backups.
Duplicity, as you may know is not a deduplicating backup solution, it is an incremental backup solution, suffering from all the flaws and limitations thereof. If you happened to not realize the difference, I highly recommend this short but informative blog post from Backblaze: https://www.backblaze.com/blog/backing-linux-backblaze-b2-du...
I know the difference, but this way we have 4G volumes we could eventually burn to DVDs if we'd want to. No backup repositories, cloud storage solutions or new and obscure network protocols involved. Deduplication is also uninteresting for us because the largest part of our backup are SQL dumps. This is why we make full backups every time.
Also why would deduplicating be uninteresting when the backups are SQL dumps? That is EXACTLY the scenario in which deduplication would shine, because it would take each dump and find parts of it that are the same as the latest backup, finding the individual small chunks that have changed in those dumps, and only backing them up. Do you actually know how deduplication works? You might want to look that up more properly.
Until one of the statically compiled dependencies needs a security fix.
Then people realize why shared objects where invented.
I'm more concerned about the repository format and config file (i.e. attack surface, since the repo is potentially untrusted).
Performance is actually better than Restic, and performance-critical parts of Borg are written in C or use C libraries.
No, using one key per repository and a persistent message counter is not a reasonable design.
https://borgbackup.readthedocs.io/en/stable/internals/securi...
Edit: I've posted this a bunch of times here, pretty much every time it caught my eye when someone said this tool has good crypto, and by now I'm used to people just downvoting it and saying it doesn't matter because obviously no one ever would use it like that and the design is fine etc. (isn't the point of deduplication to save disk space?)
> If we perform similar backups to the same remote destination over SSH (Borg) and SFTP (Restic), the initial backups take roughly the same time, but the subsequent incremental backups take something like 10x longer with Restic.
Also I'm wary of encrypting personal backups. I think the chances of me forgetting a password that I used literally once are quite high, and what's the point of a backup if you can't actually access it?
If you don't want to run your own server, I offer a hosted Borg Backup service: https://www.borgbase.com
This would give you some Borg-specific features, you can't easily get with your own VPS, like monitoring for outdated backups and restricting certain keys to append-only mode.
Firstly, if you want to prune old backups, e.g keep the last N1 hourly backups, and the last N2 weekly backups, etc, then it has that ability, however whilst it's doing it, the client has to download and upload a tonne of data in order to repackage the backup files that contain some data that needs removing and other data which doesn't.
Secondly, I've set up an "append only" system, where my various hosts can append to their own backups, but not overwrite or delete them. I wanted the backup server to be unable to read the backups (easy enough, don't supply the encryption keys to the backup server), however at the same time I wanted the backup server to be able to automatically prune old backups. It can not do that without the key. I don't want to give it the keys to the backups as then a compromised backup server means all of my hosts data are suddenly compromised.
For example, although drives in Google Cloud are encrypted at rest, once they are decommissioned the drives are physically destroyed: https://cloud.google.com/security/deletion/#ensuring_safe_an...
Thanks. I guess the reason is to avoid future decryption due to either a new crypto attack or to avoid being vulnerable to potential errors in key generation.
I believe immutable trees would be the correct algorithm here.
The algorithm relies on a tree structure, where the leaves are the files. Every time a new snapshot is written, every leaf node changed writes a new path to the root node, while leaving the rest of the tree intact.
When a snapshot needs to be deleted, you just remove all the old root nodes and have a garbage-collection like-process delete all the old files.
EDIT: Look at example here: https://en.wikipedia.org/wiki/Persistent_data_structure#Tree... . When xs is deleted, nodes a,c,f can also be deleted.
EDIT2: I feel like I can never make posts like these while keeping it short and sweet, and adding enough details. ZFS can be encrypted, but accessing any files is right at its fingertips, not over a slow network connection. And there are caveats about how to encrypt the metadata, but you do not need to decrypt the whole backup to figure out what to do delete.
ETA: However, nearly all data in restic is encrypted. This includes the index files. So you still need to have the encryption key to look at snapshots and walk their trees.
It certainly can be (as of zfzol 0.8, and oracle zfs has its own encryption scheme).
Neither.
But then, I assume the way it is written, there are some blobs that belong to multiple snapshots, and some bits of blobs that belong to one snapshot but not another etc.
So given the implementation, it may not be possible, but it's still a drawback of this particular implementation that people should be aware of.
To my knowledge that's not possible with restic. You can either have append only (using for example restic/rest-server) or pruning. But not both.
My solution so far is to use a backend like rsync.net. The backup uses the normal backup + prune routine. And rsync on their server side creates daily undeleteable snapshots of all files and deduplicates the encrypted files in the process without the need to look into them. For me that's seems like a reasonable compromise for now.
I posted this in another comment, but I did this as follows (I'm the author) - https://www.grepular.com/Nginx_Restic_Backend
$ pg_dump dbname | restic backup --stdin --stdin-filename dbname.dump --tag dbname -q
With restoring as simple as:
$ restic dump latest --tag dbname dbname.dump | psql dbname
I should probably stick a gzip/gunzip between those two commands.
If it does compression, then it's not an opaque process to it. But if you do it, you are essentially obfuscating the structure of the data.
The result is a valid gzip file with a few percent less compression, but friendly to deduplication (which is what rsync does).
Unfortunately it does not seem like it's a prioritised feature given that it has been planned since late 2014.
You cannot stick a compression step before restic, because it will mess up the deduplication - the chunker will not find the same chunks, and will not be able to find out which files have same contents (not as easily at least).
And you cannot stick it after restic, compressing the backup archives (or storing them on a compressing filesystem), because those are encrypted, and any encrypted file has high entropy and does not compress. (It doesn't help that restic makes the encrypotion mandatory).
So no, compression really needs to be an integral step in a backup system. Restic is definitely lacking in this regard.
The biggest missing feature of Restic is no support for compression.
Ah, well.
I do concede it is annoying having to fold in upstream PRs to my own build of it.
More generally, I've been looking for a solution that helps distribute backups in a peer-to-peer way. I have a few friends with their own home servers, and we want to replicate backups across each other's servers for geographical redundancy. Currently, I have a script that uses rsync to copy some tar archives over daily, but this doesn't scale well as more peers want to join our backup-sharing group, since it requires them granting me SSH access.
What I need is a decentralized network to share and retrieve backups from peers. I tried using dat [2] with a Borg Backup repository inside it, but ran into some nasty issues with dat which would cause it to regularly crash and one time even corrupt the data.
Does anyone have any suggestions for such a situation?
[1] https://www.borgbackup.org/ [2] https://github.com/datproject/dat
My plan was to use gRPC + go for communications, a DHT to find peers, and github's klauspost/reedsolomon for adding redundancy.
I wanted to support a simple client in go, tracking all filesystem state locally in something like sqlite and of course encrypting before upload. The local state of course would be backed up as well.
Encrypted blobs would be offered for upload to a peer 2 peer server (or in small setups it could be the same machine) and accepted if they were unique. If not unique, the client would be subscribed to that blob.
The server would then chunk up 1GB or so of blobs, run the reedsolomon to add the desired level of redundancy and start trading those chunks with peers. No peer would know if you trusted them, you might well set your server to only "trust" peers until they have a 95% uptime and 95% reliability when challenged over a month. Reputation would be tracked for the peers you trade with, but only directly. Much like torrent's tit for tat strategy.
The p2p server would accept uploads from any trusted clients and work to ensure the configured replication across any peers it could find.
The peer challenges would be something like ask for the sha256 of a range of bytes of a blob the peer stored for you. Maybe 100 random challenges every few hours.
The general goal is something that would "just work", create keys, get nagged to print the keys out, and have sane defaults for everything.
- I don't have to think about deduplicating my own data, which works well with my packrat tendencies (e.g., multiple copies of music library on snapshots of old laptops);
- Full-disk backups of multiple machines will share a lot of storage (all system files from my desktop and laptop, both running the same Linux distro);
- Restarting a long upload doesn't depend on some finicky state regarding an interrupted previous backup; it can just rescan the whole disk and skip uploading blobs that are already there, which feels much more robust to me.
But paid is good that the author has less reason to move away from it. (As mentioned elsewhere, restic hasn't been updated in a while but they recently changed their pricing to be much more agressive which I hated. I wish they give it a softer pricing model than cost per machine.)
Their comparison of cloud storage providers gave me a good insight on what to choose in terms of performance.
https://github.com/gilbertchen/cloud-storage-comparison/
I use both restic and Duplicacy, so my backups are done by multiple implementations toward multiple destinations not to get bitten by one of their bugs to avoid the saying "backups aren't working when you need it the most".
I've also been using Syncthing for syncing machines. It's pretty great. i've got it setup so 2 laptops sync to a server and no issues so far.
I remember reading a long time ago that Restic was not going to natively support cold storage solutions. Has anything changed?
No support for compression yet
No support for deleting data from snapshots
No support for continous backups (restic walks the directory tree for each backup).
No support for resilience from disk errors using par2 or similar
Backing up millions of small/empty files uses a lot of memory
Not too many storage backends.
No GUI tool.
Hard to implement lifecycle policies like "keep N last hourly backups and M last weekly backup"; and not optimized for this usecase (needs quite a bit of data transfer to accomplish due to the zero-knowledge server-side encryption).
The one issue I've encountered with it: it uses a lot of memory, proportional to the size of the repository indexes, so if you're backing up a lot of data on a machine without a lot of RAM (such as a virtual server), you may run out of memory. Setting GOGC=20 can help slightly, but ultimately, restic needs fixing to support working on indexes larger than memory.
Here's what I mean. I develop a software/firmware stack that is typically delivered to users as a set of large (100M) tarballs and binary images. Even though the vast majority of the payload does not change from release to release, it really has proven necessary to distribute just the final blobs.
Technologically, it's almost identical, but distribution implies a different set of access controls and an Internet-friendly user interface.
I see the idea floating around in e.g. NextCloud forums but this seems like a relatively compact problem without a obvious candidates to solve it.
I heard there's a way to have an "append only" backup or something like that. Is it possible to still prune old backups from time to time?
I personally find it reassuring that even though I might be creating and maintaining backups with sophisticated tools like borg or restic or duplicity or rclone ... at any time, and from any system I can grab those backups with dumb old SFTP.
[1] https://www.rsync.net/products/sftp.html
[2] https://forum.restic.net/t/restic-commands-for-rsync-net/216...
* I use Ansible to configure a vanilla Ubuntu 18.04 into my workstation [1]
* I keep everything that is "source code-y" in some Git repo
* I keep non-source-code-y stuff in Dropbox or Seafile (both have a restore to previous version)
I prefer everything else to be lost (e.g. some AWS, K8s credentials).
I wipe my laptop between customer projects and it works great.
[1] https://github.com/cristiklein/stateless-workstation-config
So if some of your git repos are only hosted by a company and you have no local clones, that's a bad position to be in if the company terminates your account for whatever reason they might decide. But if you have local clones, it's fine. An ansible script can easily cover this (clone every repo you have).
If there are some files only in Dropbox, and you only have 1 computer sync'ing with Dropbox at a given time, all it takes is for Dropbox to screw something up and those files are gone. I wouldn't personally be OK with this. You might not care that much about what's in Dropbox though.
Beyond that, there are some files I don't want any service to have unless they are encrypted locally first. So I don't use Dropbox for that. Anything like that I keep in my regular documents folder. But those are mostly for my main "personal" PC, and don't need to sync those to things like laptops.
Note that recovering files in syncing services tends to have a limited time. So hopefully you notice before that time runs out. I have a friend that lost his military service documents while still using a file syncing service. We assume he accidentally deleted them years ago. Couldn't recover it.
You should probably be backing up monthly or quarterly so that you have a local copy of all those repositories.
A multi-layer backup strategy with local snapshots, TimeMachine, cloud backup seems like a more sensible approach.
What is the difference between running the Dropbox client or any other backup client?
Restic or Borg has fewer sharp edges than Dropbox does on Linux and is generally going to be lesser maintenance.
A stateless system is one that does not keep any (persistent) state.
I never backup workstations or home directories, as I can always generate them again using NixOS and home manager. For source code, nothing beats git repos.
I've been keeping encrypted remote backups for ~300 Gb worth of data, which occupies ~150 Gb. On Backblaze B2, it's costing me ~$10CAD/mo, which is a _lot_ cheaper than Tarsnap.
I might try Restic (always good to have a backup backup system), but I'm not sure how ergonomic it'd be on Windows.
The web interface is what keeps me on it.
Now I just made my home folder a shared folder on Syncthing, in send only mode, so it constantly backs up to my Nas. Much faster and the backup is directly readable.
Syncthing in send-only is not a sufficient back-up strategy for 90% of use-cases, however. Without incremental back-ups, you're still susceptible to ransomware and data-loss should syncthing update after the event. I do, however, like to keep a "verbatim" copy in addition to my incrementals since they are less likely to have a "restoration" problem.
At work I have a systemd service which runs every 10 minutes and use a nfs repository. Performance is good. It has saved me once already after a botched Ubuntu upgrade.
At home I have a identical service but it runs every 6 hour and use a backblaze b2 repository. Performance is not great. However I've been backing up ~20 GB for over a year now and it has cost me less than $2 _in total_ so I'd say it's worth it.
Documented it a bit here, mostly so I can go back and copy / paste commands if I have to:
https://blog.notmyhostna.me/backup-cloud-server-to-synology-...
Does anyone know of an open-source solution that acts just like Dropbox/GDrive where it detects for any changes in a specified folder and then once detected it automatically uploads to an S3 folder?
https://forum.restic.net/t/restic-and-s3-glacier-deep-archiv...
I've largly decided on Duplicacy over Restic, Borg or Duplicati.
The only issue I encountered was with trying to back up 2.5TB of new data over a shitty German internet connection, which first took forever (restarting 20 times because they force your IP address to change once a day), and then when trying to prune the backup, ran out of RAM memory and actually corrupted it.
I'll continue to use it, just not repeating those specific steps, but it was not immediately apparent that the backup was completely unusable. Always test your backups (regardless of whether it's about restic, borg, examplecorp expert enterprise backup, or something else)!
Any news on this?
One small issue which I need to figure out is the hosts directory being empty when using “restic mount”: https://github.com/restic/restic/issues/1869
that said, i do use it personally for private and small stuff.
So does it have a GUI Version?
Since if I understand correctly both provide backup in addition to other functionalities?