Restic – Backups Done Right
restic.net
restic.net
_backup_prepare () {
export $(sudo cat <protected_credentials_file> | xargs)
}
_backup_remove_old_snapshots () {
restic forget -r <repo_name> --keep-weekly 10
}
_backup_verify () {
restic check -r <repo_name>
}
backup () {
echo "---------------Scheduled backup time---------------"
echo ""
_backup_prepare
restic -r <repo_name> --verbose --exclude="$HOME/snap" --exclude="$HOME/Android" --exclude="$HOME/.android" --exclude="$HOME/ApkProjects" backup ~/
echo ""
echo "---------------Backup done, removing old snapshots---------------"
echo ""
_backup_remove_old_snapshots
echo ""
echo "---------------Old snapshots removed, verifying the data is restorable---------------"
echo ""
_backup_verify
echo ""
echo "---------------Backup done and verified!---------------"
}
Then in cron I just have this entry: 20 \* \* \* DISPLAY=:0 kitty -- /bin/zsh --login -c 'source ~/.zshrc; backup; read'
So every time my PC is on @ 20:00, a shell window will pop-up, asking me for password and runs the backup :). Since they are incremental, it takes maybe 10-15 minutes top.restic -r <repo_name> --verbose --exclude="$HOME/snap" --exclude="$HOME/Android" --exclude="$HOME/.android" --exclude="$HOME/ApkProjects" backup ~/
To me if it requires complicated shell commands it is not really "done right".
The snippet GP mentioned includes only shell functions, invocations of restic, and echo commands. How could it be simpler?
Simple for someone who is comfortable with the terminal. Scary for those who are oyherwise familiar with GUIs only.
It makes it hard for me to recommend backup tools to the non-technical people in my life, because they're looking for GUI solutions with the corners rounded off, and I want something crunchy and scriptable.
GUIs for simple stuff and 99% of the population, CLIs for hackermans who need the advanced stuff.
A lot of backup services (tarsnap comes to mind) prioritize scriptability, as I imagine many people run backups from cron or a systemd timer.
restic -r reponame /path-to-files-to-be-backuped
Everything else in the mentioned script is optional stuff for scheduling, removing old backups, etc.
# tar cvf /dev/st0 .
Relevant manual page in case anyone's interested: https://restic.readthedocs.io/en/latest/045_working_with_rep...
export $(sudo cat <protected_credentials_file> | xargs)
Is that the same as sudo cat f | xargs export?If so, I think an important difference is that mine won't hide the exit code, and you can then answer a sibling commenter (on forgetting without checking) with 'I omitted set -e at the top'.
It looks like in this example it's used simply to trim whitespace. E.g.
$ echo ' hello ' | xargs
helloRestic can use rclone to access cloud providers it can't access and that's something we worked together on.
I use rclone directly for backing up media and files which don't change much, but restic nails that incremental backup at the cost of no longer mapping one file to one object on the storage which is where rclone shines.
(rclone author)
Is there any interaction between these two layers of encryption, or are they completely independent?
Also, I have a crypt set up for other stuff and rather deal with one remote.
Otherwise, cascades encryption is generally not recommended.
I've been using it for a few years and was not aware of it so would be great to know if any backups were affected.
Saving a restic backup the hard way - https://news.ycombinator.com/item?id=28438430 - Sept 2021 (2 comments)
Restic Cryptography (2017) - https://news.ycombinator.com/item?id=27471549 - June 2021 (5 comments)
Backups with restic and systemd and prometheus export - https://news.ycombinator.com/item?id=27411713 - June 2021 (1 comment)
Restic – Backups Done Right - https://news.ycombinator.com/item?id=21410833 - Oct 2019 (177 comments)
Append-only backups with restic and rclone - https://news.ycombinator.com/item?id=19347188 - March 2019 (42 comments)
Restic Cryptography - https://news.ycombinator.com/item?id=15131310 - Aug 2017 (36 comments)
A performance comparison of Duplicacy, restic, Attic, and duplicity - https://news.ycombinator.com/item?id=14796936 - July 2017 (49 comments)
Restic – Backups done right - https://news.ycombinator.com/item?id=10135430 - Aug 2015 (1 comment)
Edit: interestingly the restic repo's Readme is better in that regard. It's the demo video but in text form with explanations: https://github.com/restic/restic#quick-start
repo=/mnt/mybackupdrive/somefolder
restic init -r $repo
restic backup /my/files -r $repo
restic mount /mnt/backedup -r $repo
cp -r /mnt/backedup/snapshots/latest/accidentally-removed-directory ~/restored
Where the repository can also be sftp or various other things.There are a ton of other commands and options, but to literally just get started, this is all you need for making and restoring data from a backup.
This quickstart assumes some context that not every person starting to use Restic has. It could be improved by offering some more context for each line:
# RESTIC QUICKSTART FOR LINUX
# Choose were you want to save your backups
MYREPO=/mnt/mybackupdrive/somefolder
# Initialize a new restic repo at your chosen backup location
restic init -r $MYREPO
# Backup your files to the newly created repo
restic backup /my/files -r $MYREPO
# To restore a backup, first mount the repo
restic mount /mnt/backedup -r $MYREPO
# Browse the latest backup at /mnt/backedup/snapshots/latest
ls -la /mnt/backedup/snapshots/latest
# Copy the files you want to restore out of the repo
cp -r /mnt/backedup/snapshots/latest/directory-to-restore ~/restoredQuick start != guide, to me. A guide will guide you through thing at a slow pace, a quick start is the quickest way to get something running/started. And the commands seem fairly self-explanatory, but I could be suffering from the so-called curse of knowledge.
Fwiw, I've used restic for a few years, and (still, if perhaps less) agree the docs are not great.
Mostly I just use restic --help or restic subcommand --help anyhow, rather than the docs. Perhaps I'm just not doing fancy enough things with it and that's why I didn't run into edge cases yet where it lacks documentation?
Like...maybe it's time to stop working on new features, and focus on a release.
Duplicati has a similar problem. They're endlessly tinkering with new features instead of squashing bugs and meeting user expectations. I think it still chokes and permanently corrupts backup archives if you interrupt it during the initial backup.
Like: guys. People expect backup software to be able to handle interruptions and disconnects. It's actually one of the things I liked the most about Crashplan, and it could handle being interrupted without issue...nearly a decade ago.
I then use FolderSync[3] to SFTP synchronise those two backup folders across to my server regularly when the phone is on the home wifi. (I also two-way sync my photos folder which is really quite handy.)
I use to also occasionally do a full sync of my phone contents to my server using FTP[4] although since upgrading, Android 11 has clobbered access to the Android/data folder making that problematic.
Using Termux + Borg (or Restic) so push full full backups looks attractive. Never seen Termus before. Thanks.
[1] https://play.google.com/store/apps/details?id=com.keramidas....
[2] https://github.com/machiav3lli/oandbackupx
[3] https://play.google.com/store/apps/details?id=dk.tacit.andro...
[4] https://play.google.com/store/apps/details?id=com.theolivetr...
It is a shame, it has been a mainstay for me, restoring apps and data across at least three phones now. I'm hoping OAndBackupX works out, but have not really battle-tested it yet.
It is possible to do pull backups with borg, with some gruesome ssh hackery.
On the backup client side, you need to have a /root/.authorized_keys line like this (edit borg options to suit):
command="BORG_PASSPHRASE=$(cat /root/.borg-passphrase) borg create --rsh 'ssh -o \"ProxyCommand socat - UNIX-CLIENT:/root/.socket/borg-socket\" ' --compression auto,zstd -s -x ssh://borg-backup@localhost/<repo location>::<backup name> <backup dirs>" ssh-ed25519 <keydata> <keyname>
Then, on your borg server, you need an authorized_keys file like this: command="borg serve --append-only --restrict-to-repo <repo location>",restrict ssh-ed25519 <keydata> <keyname>
Finally, run a small shell script like this from the borg server to trigger a pull backup: #!/bin/bash
eval $(ssh-agent) > /dev/null
ssh-add -q /home/borg-backup/.ssh/<keyname>
ssh -A -R /root/.socket/borg-socket:localhost:22 -i .ssh/<keyname> root@borg-client
ssh-agent -k
I trigger this using cron every night, but systemd timers will work too.The first neat thing about this setup is that the client never even sees the private key that it uses to authenticate to the borg server - the key stays on the server & authentication is tunnelled between client & server via ssh-agent. You don't even need to be able to make a tcp connection from the client to the server - so long as the borg server can make an outgoing tcp connection to the client then everything just works. The client connects back to the server via a socat connection through a unix socket created by the outgoing ssh connection that tunnels any tcp connection made through it back to the sshd on the server. (You could probably tunnel the repo passphrase through as well, if you really wanted to.)
The second neat thing is the use of authorized_keys commands which are tied to an ssh keypair means that you're giving the minimal possible access - each ssh connection can only trigger that specific command & no other. You can issue ssh keys on a per-host basis & revoke them individually if necessary.
You have to use socat as a proxy program for the return ssh connection as ssh doesn't know how to connect to a unix socket & this setup requires
config
StreamLocalBindUnlink yes
in the .ssh/config on the client (possibly both client + server?), as otherwise the unix socket doesn't get cleaned up after the connection ends & the whole thing only works once before you have to remove the socket by hand. I'm not sure why this isn't the default for ssh to be honest.This method is outlined in the ssh-agent section of https://borgbackup.readthedocs.io/en/stable/deployment/pull-... but the docs don't really call it out as a method of getting pull backups working properly. It's a bit convoluted, but it does work!
(If your client can make a direct tcp connection to the server you can skip the whole song + dance with the unix sockets of course.)
It’s the unix socket dance that introduces the gruesome hackery (imo at least!).
Note: I'm not trying to be horrible or disrespectful to the Restic devs; it's just that backup without compression is a complete show-stopper, especially if you want to use cloud storage (where storage space, storage operation, and bandwidth all eat money).
I don't know if they apply to this specific case, but I assumed that the restic developers were being extra paranoid.
During initial development (versions prior to 1.0.0), maintainers and developers will do their utmost to keep backwards compatibility and stability, although there might be breaking changes without increasing the major version. "
they are on 0.12 now
In one of the early talks at a local hackerspace, the author did also demo decrypting the data manually if something got somehow broken. The tool is just a layer on top of relatively straightforward cryptography. From what I remember (I looked at this in 2018 so forgive any errors), you'd have to write a script that iterates over the index where it says which blocks are in which files (since it's deduplicated) and decrypt it with some standard AES library. Perhaps as a security consultant this seems easier to me than it does to others, though, but it's not as if you're without hope if the tool did break, or as if you couldn't just download a previous version from the GitHub releases page.
For extra paranoid, could clone the restic source tree (with vendored dependencies). Go language backwards compatibility is such that I should always be able to read my data.
I've noted this before that my NAS has a tendency to kill a lot of backup applications that haven't really been stressed. Its about 50 something TB of data, made up of a fair number of compressed video files, family movie kinds of things and about ~7T of source files. The combination of a crapton of tiny files and 50T of data kills most of these more recent open source backup applications, which seem to be released with the "it can backup my laptop" testing.
Also, I see someone mentions it sill doesn't have a compressed block option, which I really want because most of those source files compress to about 1/3rd of their space, which is important since my upload speeds are crummy (thanks spectrum!) I wonder how much of that deficiency is just lack of good go compression routines that can run at a few hundred MB/sec.
I'd definitely recommend trying restic again if it's been a couple years. Somewhat recently they made some nice speed improvements. It used to take me several days to do a forget+purge on my restic repo. Now it takes only a few hours (less than an actual backup takes).
How many files are you backing up? I found the biggest issue for me was that almost every tool tries to keep the list of files in memory, so once you get into millions of files - it starts to require a lot more memory and can crash on low-resource machines.
However for desktop use, I've always really struggled with the idea of not having a UI for my backup client. I'm not afraid of the command line, but the idea of browsing backup archives without a GUI feels awkward to me.
I wonder if there is room for some sort of add-on GUI for Restic for those that are more visual (unless such a thing already exists?)
I think many more people would use Restic if this was called out prominently on the home page.
I've known about Restic for years and would likely be using it by now had I realised that!
EDIT: Ah looks like it may not work on Windows. Part of the appeal of Restic, for me, would be being able to cover all of my Windows, Mac and Linux machines with the same system.
I do like the idea of being able to tune and configure things from the CLI. It was frustrating configuring CrashPlan on a remote computer.
With that said, I feel like both should be possible. Even just a basic wrapper GUI would be a start.
Edit: and some basic searching has lead me to lots of options! Time to do some more research.
However I also acutely feel the lack of a standalone GUI so that I can get rid of the scrips and custom setup or at least that can be an option. (There’s a commercial third party UI I think which is a subscription)
https://vorta.borgbase.com (a third party qt GUI for borg backup) has been really awesome.
Another tool I’m looking at closely is https://kopia.io. It comes with a UI by default (Electron I guess). Though its UI and logo has quite some work left.
Related: "Restic Cryptography" by FiloSottile (2017), 36 comments https://news.ycombinator.com/item?id=15131310
I usually work in multiple Linux virtual machines and I have a Bash script to setup regular backup of all my many $HOMEs.
I did not yet have the script for verification (did it manually just to be sure), but the rest is here if someone is interested https://github.com/senotrusov/sopkafile/blob/main/lib/ubuntu...
I pay around 0.40$ a month to have a backup of around 70 GB.
Also Borg has an append only mode.
Note that write-only or append-only is a bit of a misnomer, since reading the files is fine (you need the decryption key anyhow before they're of any use). It's about not being able to overwrite or remove backup data without some verification that you're not ransomware or similar.
You could do deduplication on the encrypted blocks I suppose (is that secure?)
That's always an option, no matter if it's set to append-only or not. If you don't want to pay infinite costs for storage, you will need to limit it using software on the server.
As a proof of concept, take a look at a Certificate Transparency log server. Most CT logs are configured to only accept certificates meeting certain criteria. They'll log any such certificates (their SLAs only apply to contractual users, but you don't need an SLA you're probably just writing one certificate to the log to see it works) but you cannot fill them with garbage because you can't make any certificates they'd accept, only the legitimate CAs can do that†
† The CAs have their own reasons not to let you produce heaps of garbage, even Let's Encrypt has finite resources and so it imposes rate limits.
Given the chance, there certainly will be people that use 2FA when making an append-only regular backup, but even among command line restic users I expect this will be the exception rather than the rule.
That said, Tarsnap isn't free software, and doesn't come with the ability for me to self-host backups. Thus, I am somewhat at the mercy of Amazon.
Still, the advantages of Tarsnap currently outweigh the disadvantages in my opinion.
For the record, this is not a shill. I don't even know who Jenny is.
Quite happy with it.
Other than performance, the biggest benefit is that multiple clients can write to one repository at once. There is a designated maintenance user, though.
From what I learnt, borg has least problem of all the open source backup software with the most wanted features (encryption, dedup, compression and rotation) and others all have their quirks, be it huge memory usage, data loss to worst being corruption of the backup repository.
Things may have improved since I've looked but you want user feedbacks saying so.
It's too late to know that backup software has been failing when you need it as the original data is already unavailable, so choosing by the look of landing page is a bad way when it comes to backup.
kopia is still new and last time I used about a year ago, it still had basic problems, so I would never use it in place of borg.
restic, duplicati, duplicity and duplicacy all have some sort of problems especially when the repo gets large but of course there are cases where things are working fine.
https://forum.rclone.org/t/rclone-as-destination-for-borgbac... (Restic issues)
https://www.reddit.com/r/Backup/comments/opu1ep/comment/h67n... (Kopia issues)
https://forum.duplicacy.com/t/memory-usage/623/24 (Duplicacy doing a big rewrite of the engine to fix memory issues.)
https://forum.duplicati.com/t/is-duplicati-2-ready-for-produ... (duplicati stability issues)
https://www.reddit.com/r/unRAID/comments/eg0zpe/duplicati_se... (Another duplicati issues)
borg also has change logs with any major reliability problems mentioned up front which gives you more confidence than serious bugs buried in GitHub issues like other tools.
https://borgbackup.readthedocs.io/en/stable/changes.html
The only downside of borg is it can only target ssh host natively, but there are services like rsync.net (with special borg pricing), borgbase or you could locally run borg and rclone the entirety to anywhere you want.
Every software has its issues. That's why it's important to test restores regularly.
Slowness is the least of the problem against data loss or repo corruption and perhaps you may be able to somewhat circumvent it by splitting the repo to different locations or disks.
https://duplicity.gitlab.io/duplicity-web/vers8/duplicity.1....
> When restoring, duplicity applies patches in order, so deleting, for instance, a full backup set may make related incremental backup sets unusable.
Other software mentioned have each incremental backups independent of each others, so you can prune any of it and the last one is still retrievable.
Just mind you exclude node_modules and the like from it :)
Anyway! I'm happier with restic now. It's never crashed for me, and it has native cloud backends. But it's ultimately just another backup application.
I use it over NFS.
To which the conclusion is: test your backups!
Still, some might corrupt more easily than others. People just confuse first-hand problems with statistical significance.
Restic scans your filesystem for changes, and then also starts backing up the changes it finds in parallel while it is still scanning for more changes.
When you have millions of files, this makes a huge difference.
With restic, it still took around 10 hours to scan for changes, but it was also already done backing up all the data by the time the scan finished.
attic (python) - https://github.com/jborg/attic
borg (c) - https://github.com/borgbackup/borg
bupstash (rust) - https://github.com/andrewchambers/bupstash
duplicacy (go) - https://github.com/gilbertchen/duplicacy
duplicati (c#) - https://github.com/duplicati/duplicati
duplicity (python) - https://github.com/henrysher/duplicity
kopia (go) - https://github.com/kopia/kopia
nfreezer (python) - https://github.com/josephernest/nfreezer
rdedup (rust) - https://github.com/dpc/rdedup
restic (go) - https://github.com/restic/restic
rclone (go) - https://github.com/rclone/rclone
rsnapshot (perl) - https://github.com/rsnapshot/rsnapshot
snebu (c) - https://github.com/derekp7/snebu
tarsnap (c) - https://github.com/Tarsnap/tarsnap
I think there are many more out there (https://github.com/restic/others) - I personally use
restic
while technology wise (speed, only restore needs password) i would prefer rdedup
which is an impressive piece of software but unfortunately without file iterator... :-)zpaq (C++) - http://mattmahoney.net/dc/zpaq.html
to the list.
But zpaq is pure magic...a slow one but with massive compression.
borg is python and c
bupstash is rust and c
Maybe update your list a bit ;)
I’m trying to avoid the need to reinstall and configure my system (for example, the registry, custom installed and tweaked programs) in case of complete data loss or a migration to new hardware.
I use Veeam Agent for this purpose (free, but not open source). It can do full system backups and supports both restoring to the same hardware and new hardware. Restores are done via a bootable WinPE-based image that the tool creates.
One cool thing about it I haven't seen in other backup software is that incremental backups work via a driver that tracks which disk blocks are changed as the system is running. It avoids the need to rescan the disk to detect what has been changed (though it will still do that if the filesystem is modified outside of Windows, eg. if dual booting).
The biggest downside is Veeam's website. It's pretty "enterprisey" and they want you to register to be able to download. I install via the Chocolatey package manager to avoid this. Chocolatey's package source has a direct link to the official installer [0].
There are no ads, nagging, nor upselling in the software itself. I have not seen it making any network connections outside of connecting to my backup target host and the auto-updates server.
I've been looking for open source alternative with a similar feature set, but haven't had too much luck. There's Bacula, but that seems to very much be designed for an enterprise use case.
[0] https://github.com/sbaerlocher/chocolatey.veeam-agent/blob/m...
I do miss Obnam, nevertheless. The interface was the best.
No support for compression yet
No support for deleting data from snapshots
No support for continous backups (restic walks the directory tree for each backup).
No support for resilience from disk errors using par2 or similar.
No directory chunking, so backing up millions of files in one directory uses a lot of memory.
Previously on HN:
However, in this case it's the open source tool that has a much easier user interface (I am actually proficient with tar, but still my tarsnap experience is like comparing 'restic backup /my/files --repo /mnt/backupdisk' with https://xkcd.com/1168/)
What is your definition of "service" that makes tarsnap - a company that asks you to pay them over time to provide an, uh, service - not one?
For years we had a running discussion here, where patio11 writes a nice article called something like:
fake patio11> "10 reasons why tarsnap must raise the price and stop using funny units"
And a week later cperciva writes another article called something like:
fake cperviva> "Nah. Amazon reduced the storage price 10%, so I'm reducing the price of tarsnap in 1 picodollar/byte"
I would be sad if either of those ceased.
Example, if you are into photography it's not uncommon to generate hundreds of GBs of files _per year_. Only in 2020, I generated over 200GB of photographs. Putting that on Tarsnap would cost me about $60/month. In 5 years time, I could be paying upto $4000/yr. Tiered services like B2 would cost an order of magnitude less.
Raw managed storage with rsync.net is 0.18$/GB/year.
Do it all yourself, with the associated peril and time sink that entails, and the disks will cost you 0.04$/GB/replica.
Tarsnap has its place and I’m still a happy customer, but it’s one small part of a wider strategy that includes bulk storage elsewhere — rsync.net with borgbackup and plain rsync, on premises ZFS dumpsters, and offsite drives used like they are tapes.
i have 700G so it would cost me ~ $230/year (and you have to buy a min of 400G)
microsoft onedrive is $70 for 1TB and you don't pay for bandwidth. can use rclone.
you get some other goodies too (office) which i don't use but i'd imagine it's a nice to have for some people.
I’ve also never seen successful backups anyplace that did not have TSM. Usually the backups are corrupt, or nobody knows how to do a restore, or you need 1000x the storage capacity in order to restore the initial backup and all of the incremental backups until you reach a specific point in time.
At places I worked with TSM it was so simple that individual users could fire up a gui and pull files out of the backup pool.
On the backend we had massive IBM tape libraries and it was hypnotic to watch the robot jet around moving tapes in and out of drives and the storage slots. It never stopped moving either, when backups are restored were not happening it was busy consolidating the tapes, making copies of data from tapes that had been used too many times, or preparing copies to be sent off site. It was a full time job for someone to load new blanks when TSM requested and remove the offsite tapes and put them into a box for fedex to pick up. (The one thing that has not changed is that it’s still quicker to send massive amounts of data by fedex then it is to send it over the public internet)
TSM also has no support for deduplication, so good luck backing up large variable binary files such as VM images or project files (video, CAD, etc).
They also have optional replication, and you can choose the location of both of the servers that host your files. Great support from actual engineers, used by huge companies for their offsite backups, and never been breached (to my knowledge).
EDIT: their features page is a good read. https://www.rsync.net/cloudstorage.html
Any cloud service is going to be 5-10 times more at least, last I checked (2019), and that's already a lot better than five years before that (as a ratio of self-hosted to a managed service, so independent of raw storage prices).
B2 and Glacier seem to be some of the cheaper options these days. Backblaze (the backup software) doesn't run on Linux and is closed source but they pinky promise to support "unlimited" backups for $5/month which is a really good deal if you both trust them and run Windows. Tarsnap is S3 pricing plus markup, but what you get is linux support, open source clients from a person the community trusts, write-only keys, and the hosting part is off your hands, and pay-as-you-go, which is quite a unique combination.
Edit: B2 is $5.7/TB/month according to https://news.ycombinator.com/item?id=29209665 (`0.4/70*1000`)
About the "raspberry pi" thing... this is kind of answer you immediately regret you didn't preemptively dismiss in the question. I mean, it's hard to even decide if the person saying this is serious or not. Like, setting up a backup server at "your friends house"? Really? Is this seriously something that everybody but me does? Should I do that too? Is it considered normal practice in their cultures? Or is it just something that they say, because they like giving advice they don't follow? To me, that sounds just crazy.
Paying under $100/year to know that all of my junk is safely kept somewhere doesn't sound crazy at all, on the other hand. But which is the best option and if they really keep their promises I don't know, of course, that's exactly why I'm asking. Maybe there's some problem in disguise, maybe hardly anybody even uses their services and I shouldn't trust them. I don't know.
I never said it did, unless you are confusing Backblaze (the name of their backup solution, https://www.backblaze.com/cloud-backup.html) with their separate and much newer B2 service.
> this is kind of answer you immediately regret you didn't preemptively dismiss in the question
Well I'm sorry.
For what it's worth, this is literally the solution I use so I thought I could be helpful by at least including that as a base price point.
> which is the best option and if they really keep their promises I don't know, of course, that's exactly why I'm asking.
I don't understand, has anyone shared stories of paying Google/Amazon/Microsoft/Backblaze/Dropbox for X amount of storage and them not keeping their end of the contract? I understand your question even less now than I thought I did before.
Take B2 for example. I know their storage pricing (that's very easy to find on their website), and I know for a fact it's super affordable, compared to other similar services. But I also know, that the "fine print" in their case is just that the upload speed is the bottleneck, which will prevent most users from backing up too much. And the fact that they have only 4 facilities. Is the latter the real problem? Well, I don't know. I didn't hear any stories about them loosing user data, but that might be just it — I didn't hear them. That's why asking such questions on places like HN has value in my opinion.
Similarly with HDD cost. It's kinda obvious that this is the most affordable solution, so if the person asking doesn't know that, it means he didn't do his homework. I don't use my friends' houses for that (that really sound super awkward), but in fact something similar is my current "solution" as well. But it feels like something I should be adviced to stop rather than to start doing. Backup service backend needs maintenance too. And maintaining it doesn't seem like a fun hobby (that's part of a reason why asking a friend to do this for you seems very weird). Backups are something you want to be reliable, by definition. And HDDs are not.
(As a side-note, I also sometimes contemplate if backing up rarely changing info to more reliable storage, like tapes and optic drives is a viable option. It still seems like no, unfortunately. But keeping everything on HDDs that I personally own makes me feel uneasy as hell. Despite being something I do, this is basically a strategy equivalent to "just hoping that everything will be ok". I don't do any real work, to make sure it's the case. I have no idea, when they are likely to fail. I just hope that they don't get broken at the same time, and that's it. I have no idea what the actual probability of that is.)
I like to leave an offline backup in my car, too
> During initial development (versions prior to 1.0.0), maintainers and developers will do their utmost to keep backwards compatibility and stability, although there might be breaking changes without increasing the major version.
Hmm not sure I would like to try on a backup tool that might introduce breaking changes that easily
That said, I keep ~/Documents and ~/dev these days in Syncthing directories, and one of my syncthing nodes is an Ubuntu LTS server with zfs, with the zfs-auto-snapshot package installed.
I still run my old backup system periodically (once or twice a month) but I now think Syncthing is at a point of reliability where realtime cross-machine sync is now my primary safety net wrt "the machine in front of me has turned to entropy", versus some point-in-time backup.
> Is Syncthing my ideal backup application?
> No. Syncthing is not a great backup application because all changes to your files (modifications, deletions, etc.) will be propagated to all your devices. You can enable versioning, but we encourage you to use other tools to keep your data safe from your (or our) mistakes.
I'm really not super familiar with Syncthing, but it sounds like it's niche is availability as opposed to durability.
Anyone taking backups seriously already knows the 3-2-1 rule anyway.
This is the curse of sharp tools. Don't run unix if you don't want rm to rm.
I still also have backups, and zfs snapshots on a different independent machine.
I'll have two of the same RPi servers in different locations with all software running (like a hot swap) with Syncthing keeping them in sync while ZFS snapshots keeps a history. I can plug in a cold storage drive in the second dock slot once in a while too. I'll have a spare RPi4b on the shelf in case it dies. If my server dies I can take the off site hot backup home and reconfig the network and then it is my primary server. With remote duplicacy backup I'm days away from getting going again. So ZFS snapshots + Syncthing and cold storage is where I'm going (for home use). Also I want to stick with Linux because I set up Freenas 5 years ago and now I forget how to admin it so I'd rather just keep with ZFS and Linux (the zfs send from freenas to zfs recv on Linux works perfectly).
Here's an old discussion: https://news.ycombinator.com/item?id=19485783
That's how I use it (with --one-file-system), but I haven't tried restoring that. I just check that my data is there (e.g. open some pictures) and call it a day. For me, if I need to reinstall the system anyhow (super rare), I'm also happy to just reinstall the OS, do a bunch of apt installs, restore a few files in /etc perhaps, and put my homedir back. So I can't vouch for whether it will store all special attributes on system files, but I would expect that common things like owner/group/mode are there.
Fwiw, I once rsync'd a remote root partition to the current machine and that worked. Would not recommend to depend on this, but it apparently doesn't take much to be able to restore the root and boot from it :). You can also quite easily test this in a VM, if you'd like to make sure it works.
export \
B2_ACCOUNT_ID=123456 \
B2_ACCOUNT_KEY=DEADBEEF\
RESTIC_REPOSITORY=b2:myname-restic-myhost \
RESTIC_PASSWORD=CAFEBABE
restic backup --verbose --host myhost \
--exclude /not_this \
/yes_this \
/and_thisFurthermore, Duplicity is written in Python so you get tracebacks and quick fixes are really easy to do.
How does Restic compares to it?
30 4 * * * /usr/bin/duply myjob backup_purge_purgeFull --force > /tmp/duply_myjob.log
Does restic do this?
$ restic mount --repo /mnt/backupdrive/myrepo /mnt/backedup
$ echo /mnt/backedup/snapshots/*/path/to/file
And it would print which snapshots it is contained in.If you don't know the path, I don't know how long `find -type f` takes so you might be right about that being inefficient. It is certainly not a use-case that I think restic ever had in mind.
I also don't know of "backup" software to solve this, it seems a bit out of scope for most of them (they are meant to have a backup copy of your disk, not be a file manager with history). You might have more luck with tools like rsync (or Toucan, from the good old days where I used Windows and Portable Apps), those can certainly create only new files and never delete deleted files, but then you typically don't get encryption and deduplication.
I haven’t yet moved to restic, so it might be that find -f is good enough for me.
Yes, a backup tool never deletes by itself, but the standard way of using such tools is that you keep the last n snapshots - this is just delayed sync.
I don't know anything about backup architectures, but is it possible to be more efficient with storage and indexing(for search) if your use-case includes finding files from 5 years ago, but not necessarily recreate the whole filesystem as it was on some date?
Restic, Kopia and most other modern backup solutions allow you to to define the retention policy. Usually by specifying how many hourly/daily/weekly/monthly/yearly snapshots to keep. They don't just remove all snapshots that are older than n.
They also usually let you restore any given file, from any given snapshot, without having to restore everything. And as long as there's a functional index, this shouldn't be terribly slow either.
If a file existed for 3 days and was then deleted, it might make it to 3 daily snapshots, and no weekly, monthly or yearly snapshot.
Even if you’re lucky and the snapshots are not deleted, if you don’t remember when the file existed, currently you need to search through all snapshots. There’s no index or any data structure to speed this search up.
I agree, restoring is fast once you’ve found the file.
Currently I use backblaze, and it does have infinite retention and snapshots, and restoring is fast. But I’m looking to migrate to something that gives me search as well.
Since I have a directory full of these logs, I can search for deleted files by doing something like 'grep "^- " restic-diff-*.log'
I'm not sure if there's a better way to do this, but it works pretty well for me.
I've lost a few months of data twice because of my ignorance
Has saved my bacon a few times, I don't care about incremental snapshots or delta's (though I've used both rsnapshot and rdiff successfully in the past), I'm covering the "if the SSD blows up, how long to chuck a new one in and be back up" case not the "I might need that file from 6mths ago".
I also have syncthing setup via /home/<user>/Shared/{Personal,Work}/ on every machine and important stuff I just chuck in Shared/ and forget about/it's available wherever I need it at any point.
I've had really ornate bulletproof snapshot based backups but honestly for my particular use case they where more hassle than they were worth, rsync does what I want every time and has never let me down.
A startup I worked for [0] shipped a backup/"DR" appliance using rdiff-backup behind the scenes that I had the pleasure of inheriting ownership of.
What rdiff-backup was good for was creating the false impression that you had working backups you could restore from. But once your available disk space for backups filled up, which is kind of the whole goal of a backup system; accumulate as many revisions going back as far as you have space for, the thing paints itself into a corner you can't recover from without creating potentially huge amounts of free space first.
Here's why:
1. The backup tree is modified in-place in the course of performing a backup. If the backup is prevented from finishing for any reason (admin/user cancelled, ENOSPC, power loss, backup source became unavailable, etc), the backup tree is left partially in the new revision and partially in the previous revision. Any subsequent operation, restore or backup, must first restore the backup tree to its previous version, using the same primitive restore algorithm rdiff-backup uses for general restores.
2. Unless you're restoring from the latest revision requiring no reassembly from differentials, the restore algorithm requires enough free space to store up to two additional copies of any given file having changes it's reassembling. This doesn't even include the final destination file, if restoring into the backup filesystem (as it does when recovering from an interrupted backup, mentioned in #1), you need space for the third copy too.
Maybe they've fixed these problems since my time dealing with this, it's been years.
I ended up writing a compatible replacement for my employer at the time which used hard link farms to facilitate transactional backups requiring no recovery process when interrupted. This also enabled remote replication to always have a consistent tree to copy offsite while backups were in-progress, something rdiff-backup's in-place modification interfered with. As-is you'd end up just propagating a partially updated backup offsite if it happened to overlap with an ongoing backup.
My replacement also didn't require any temporary space for reassembling arbitrary versions of files from the differentials. So it could always perform a restore, even with no free space available. I even built a FUSE interface and versioned backup fs virtualization shim for QEMU+qcow2 atop those algorithms. But it was all proprietary and some of the stuff got patented unfortunately.
I wouldn't consider rdiff-backup usable if it didn't at least have the ability to restore without free space yet. At least then it might still be able to do its rollback process when ~full, assuming it's still doing the in-place modification of the backup data.
Edit:
In case it's not clear from the above; it's particularly nefarious the way rdiff-backup would fail, since it was typically unattended automated backups that would fill the disks, leaving the backup tree in a rollback-required state to either run another backup OR restore. The customers usually discovered this situation when they urgently needed to restore something, and rdiff-backup couldn't perform any restore without first doing the rollback, which it couldn't do because there was no space available. Not that it could even perform a differential restore without free space, but the rollback-required state almost guaranteed a differential restore was required just to do the rollback.
Back when I was implementing the replacement it was such an urgent crisis that I was logging into customer appliances to manually restore files from sets of differentials without needing temporary space, using unfinished test programs before I had even started on the integration glue to streamline that process.
Personally I would use restic in your situation because I'm familiar with it already and it does what I need (I particularly like the encryption aspect), but that's not to say that borg, bup, rsync, or other tools couldn't also fit your needs.
Would there be any advantage for me to look into a self-hosted alternative like Restic?
(I use Borg and the feature is an absolute must for me)
EDIT: see "blob" https://restic.readthedocs.io/en/latest/100_references.html#... and "efficient" https://github.com/restic/restic#design-principles
If I remember correctly, another tool either does both or does compression instead of deduplication, this might have been borg or bup. In case that's something you care about.
Original size Compressed size Deduplicated size
This archive: 167.99 GB 136.78 GB 53.85 MBAll archives: 2.93 TB 2.38 TB 133.66 GB
Sorry about the formatting. But the compression is not completely irrelevant. Dedup of blocks between files and backups is of course the absolutely most crucial part though.
compress
data | dedup
1 backup: 168G 137G 54M
Σ backups: 2.9T 2.4T 134G
Please correct me if 'one backup' and 'all backups' is an incorrect interpretation of 'all archives'. I wasn't entirely sure what you mean by that but I think I get the point.So in conclusion, adding compression saves about 17% (a $10 monthly bill would be $8.28 instead, if you pay per GB) for you.
Even if I'd get the better of the two values, I don't have enough systems that a backup being 2/3rds of the original data size reduces the number of backup disks I need to buy, but it's not insignificant either.
Deduplication means ~20x smaller backups so this is absolutely key for me.
What does restic use under the hood?