The various scripts I use to back up my home computers using SSH and rsync
github.com
github.com
The only thing missing is -> I'd like to stop syncing code with Syncthing and instead build some smarter daemon. The daemon would take a manifest of repositories, each with a mapping of worktrees->branches to be actualized and fsmonitored. The daemon would auto-commit changes on those worktrees into a shadow branch and push/pull it. Ideally this could leverage (the very amazing, you must try it) `jj` for continous committing of the working copy and (in the future, with native jj formart) even handle the likely-never-to-happen conflict scenario. (I'd happily collaborate on a Rust impl and/or donate funds to one.)
Given the number of worktrees I have of some huge repos (nixpkgs, linux, etc) it would likely mark a significant reduction in CPU/disk usage given what Syncthing is having to do now to monitor/rescan as much as I'm asking it to (given it has to dumb-sync .git, syncs gitignored content, etc, etc).
Are you really hitting that much of a resource utilization issue with syncthing though? I use it on lots of small files and git repos and since it uses inotify there's not really much of a problem. I guess the worst case is switching to very different branches frequently, or committing very large (binary?) files where it may need to transfer them twice, but this hasn't been a problem in my own experience.
I'm not sure you could really do a whole lot better than syncthing by being clever, and it strikes me as a lot of effort to optimize for a specific workflow.
Edit: actually, I wonder if you could just exclude the working copies with a clever exclude list in syncthing, such that you'd ONLY grab .git so you wouldn't even need the double transfer/storage. You risk losing uncommitted work I suppose.
Thus, syncthing basically constantly has to rescan. It's not great.
And yes, rebasing linux+nixpkgs on even an hourly basis is absolutely devastating. lol
We backup our data storage for an entire HPC cluster, about 2 PiB of it to a single machine with a 4 disk shelves running ZFS with snapshots. It works very well. Simple raunchy every night, and snapshotted.
We use the backup as a sort of Time Machine should we need data from the past that we deleted in the primary. Plus, we don’t need to wait for the tapes to load or anything.. it is pretty fast and intuitive
I have something similar; it's Nextcloud + restic to AWS S3, but it's the same principle. You can give people the convenience and human-comprehensibility of sync-based sharing, but also back that up too, for the best of both worlds. Though in my case the odds of me needing "previous versions" of things approach zero and a full sync is fairly close to backup, but, even so I do have a full solution here.
(Beyond even the fact that ~/code is also on a ZFS volume that is snapshotted and replicated off-site, which I argue can be used in all of the same important ways any other "backup" is used.)
Hence the comment! After all this blockchain hoopla and everyone's understanding of how "cool" Git is, we really, really deserve better in our backup tools.
So to bolster that other thing, I just have a simple bash script that reminds me every 7 days to make a copy of that folder somewhere else on that machine. It's not precise because I often don't know what machine I will be using, but that creates a natural staggering that I figure should be sufficient of something goes weird and lose something; like I'm likely to have an old copy somewhere?
Simplest way to think about it is that a backup must be an immutable snapshot in time. Any changes and deletions which happen after that point in time will never reflect back onto the backup.
That way, any files you accidentaly delete or corrupt (or other unwanted changes, like ransomware encrypting them for you) can be recovered by going back to the backup.
Replication is very different, you intentionally want all ongoing changes to replicate to the multiple copies for availability. But it means that unwanted changes or data corruption happily replicates to all the copies so now all of them are corrupt. That's when you reach for the most recent backup.
That's why you always need to backup and you'll usually want to replicate as well.
Ideally there is also an offsite and inaccessible from the source component to this strategy. Usually this level of robustness isn't present in a "replication" setup.
Long term storage usually has some form of Forward Error Correction (FEC) protection schemes (for bitrot), and often backups are segmented which may be a mix of full and iterative, or delta backups (to mitigate cost) with corresponding offline components (for ransomware resiliency), but that too is very dependent on the environment as well as the strategy being used for data minimization.
> Usually this level of robustness isn't present in a "replication" setup.
Exactly, and thinking about replication as a backup often also gives those using it a false sense of security in any BC/DR situations.
*edit* my gitea server saves its backups to synology
https://borgbackup.readthedocs.io/en/stable/
My two 80% full 1tb laptops and 1tb desktop backup to around 300-400G after dedupe and compression. Currently have around 12tb of backups stored in that 300G.
Incremental backups run in about 5 mins even against the spinning disk's they're stored on.
I've been using Borg, Restic and Kopia for a long time and Kopia is my personal favorite - very fast, very efficient, runs in the background automatically without having to schedule a CRON or anything like that.
Only downside is that the backups are made of a HUGE number of files, so when synchronizing it can sometimes take a bit of time to check the ~5k files.
Critically, I'm specifically referring to code sync that needs to operate at a git-level to get the huge efficiencies I'm thinking of.
Syncthing, or borg, scanning 8 copies of the Linux kernel is pretty horrific compared to something doing a "git commit && git push" and "git pull --rebase" in the background (over-simplifying the shadow-branch process here for brevity.)
re: 'we deserve better' -- case in point, see Asuran - there's no real reason that sync and backup have to be distinctly different tools. Given chunking and dedupe and append-logs, we really, really deserve better in this tooling space.
But git commit doesn’t do that. If you want to do that in git, you typically do it before commit with “git add -A”.
* Mounts the external drive.
* Starts Restic's Rest Server.
* SSH into each machine to be backed up to kick off the script that will backup to the above server.
* Stop Rest Server.
* Rsync the external drive to my office (which is a 25 minute drive away) for off-site protection.
* Unmount the external drive.
* Emails the results of it all to me.
Has been working really well so far.
Notes:
* The RPi has limited SSH access to each machine. The only thing it can really do is start the backup script on the machine.
* The Linux machines are on all the time, but the Windows machine sleeps. So first it sends a wake-on-lan. Using Cygwin for SSH and scripting on Windows. The script on the Windows machine sets the power configuration to not go to sleep during the backup, and restores the setting afterwards. Restic's ability to create a VSS snapshot on Windows is awesome.
* I still need to incorporate my two kid's Windows laptops into the backup somehow. I doubt the wake-on-lan tricks will work reliably with them. I've yet to explore Urbackup which I think I can use to have them backup to my Linux desktop periodically when they are awake.
Regarding Windows:
I have successfully mirrored a notebook and a desktop[0] (single user) with Windows using robocopy, which is a utility that comes with Windows (used to be part of the Resource Kit but I think it is now in the base product). When I say "mirror" I mean I can use either machine as my current workstation without any loss of data, as long as I run the "sync" script at each switch.
I use "net use" to temporarily mount a few critical drives on the local network, then robocopy does its work, it has maybe 85% of the same functionality of rsync (which I also used extensively when administering corporate servers and workstations). Back in the DOS days, I wrote my own very simple version of the same thing using C, but when robocopy came along I was glad to stop maintaining my own effort.
[0]or two desktops, using removable high-capacity media like Iomega zip drives.
I've used the cwRsync[0] binary distribution of rsync on Windows for backups. I found it worked very well for simple file backups. I never did get around to trying to combine it with Volume Shadow Copy to make consistent backups of the registry and applications like Microsoft SQL Server. (I wouldn't expect to get a bootable restore from such a backup, though.)
for %i in (C D) do robocopy %i:\ \\backup-server\b-%COMPUTERNAME%\%i /MIR /DCOPY:T /NFL /NDL /R:0 /W:1 /XJ /XD "System Volume Informatiowsn" /XD "$RECYCLE.BIN" /XD "Windows" /XD "Windows.old"I have my /home in a separate dataset that gets snapshotted every 30 minutes. The snapshots are sent to my primary file-server, and can be picked up by any system on my network. I do a variation of this with my dotfiles similar to STOW but with quicker snapshots.
sudo zfs send -cRi db/data@2022-12-08T00-00 db/data@2022-12-09T00-00 | ssh me@backup-server "sudo zfs receive -vF db/data"
Another alternative is to create a clone from a snapshot, which also makes the data writable.
--link-dest=DIR hardlink to files in DIR when unchanged
Basically, you list your previous backup dir as the link-dest directory, and if the file hasn't changed, it will be hardlinked from the previous directory into the current directory. Pretty nice for creating time-machine style backups with one command and no SSH.Also works a treat with incremental logical backups of databases.
--link-dest is just an elegant, built-in way to create "hardlink snapshots" the same way that 'cp -al' always did.
But note:
A changed file - even the smallest of changes - breaks the link and causes you to consume (size of file) more space cascading through your snapshots. Depending on your file sizes and change frequency this can get rather expensive.
We now recommend abandoning hardlink snapshots altogether and doing a "dumb mirror" rsync to your rsync.net account - with no retention or versioning - and letting the ZFS snapshots create your retention.
As opposed to hardlink snapshots, ZFS snapshots diff on a block level, not a file level - so you can change some blocks of a file and not use (that entire file) more space. It can be much more efficient, depending on file sizes.
The other big benefit is that ZFS snapshots are immutable/read-only so if your backup source is compromised, Mallory can't wipe out all of the offsite backups too.
> We now recommend
Who's we?
However, compression is useful in proportion to how crappy the network is and how compressable the content is (e.g., text files). This repo is about backing up user files to an external SSD with high bandwidth and low latency, and applying compression likely makes the process slower.
If your workload is IO-bound, then it is quite likely that compression will help. Most people, on their personal machines, would likely see IO performance “improve” with filesystem level compression.
My personal recipe is less sophisticated:
- znapzend on all home machines send to a homeserver regularly (with enough storage), partially replicated between desktops/laptop
- homeserver backup itself via simple incremental zfs send + mbuffer with one snapshot per day (last 2 days), one per week (last 2 w) and one per month (last 1 month) offsite
- manually triggered offline local backup of the homeserver on external USB drives and a physically mirrored home server, normally on weekly basis
Nothing more, nothing less. On any major NixOS release update I rebuild one homeserver and a month or so later the second one. Desktops and homeserver custom iso are built automatically every Sunday and just left there (I know, it simply took to much time checking so...).
Essentially in case of a fault of a machine I still have data, config and ready iso for a quick reinstall. In case of logical faults (like a direct attack who compromise my data AND zfs itself) there is not much protection beside different sync times (I do NOT use all desktops/latptop at once, when they are powered off they remain behind and I have normally plenty of time to see most casual potential attacks.
Long story short for anyone: when you talk about backups talk about how you restore, or your backups will probably be just useless bits a day...
(So, less data loss, in event of a malicious intruder on the client, or some very broken code on the client that gets ahold of the SSH private key.)
It is Awesome !
It’s very fast usually I struggle with backup tools on windows clients. And it ticks all my needs. deduplication, End-to-End Encryption, incremental Snapshots with error Correction if any, mounting snapshots as a drive and using it normally or to restore specific files/folders, Caching. The only thing that could be better is the GUI but it works.
In that regard, I don't trust anything but Borg and zfs.
I understand the benefits for deduplication etc. but this is a show stopper for me. I greatly prefer to be able to navigate my backups with cd and ls or the file manager in the GUI and inspect the files directly without having to extract them first. After all I only have to backup a laptop and little else.
Rsnapshot is hard to break by using very basic principles of file system based files and hard links. If your file system isn't zfs, I think it's a viable backup strategy for local copy while you can use others to take remote backups.
I have Duplicati [0] that does a backup of the data of my many self hosted applications Every day, encrypted and stored in a folder on the server itself.
Only the password manager backup is not encrypted by Duplicati, because it's encrypted using my master password, and it stores all the encryption keys of the other backups.
Then, I have a systemd service to run rclone [1] every day after the backups finished to sync the backup folder towards :
- Backblaze B2
- AWS S3 Glacier Deep Archive
For now I only use the free tier of B2 as I have less than a GB to backup, but that's because I haven't installed next cloud yet !
However, I still like using S3 because I am paying for it (even though deep Archive is very cheap) and I'm pretty sure if something happens with my account, the fact that I'm a paying customer will prevent AWS from unilaterally removing my data (I have seen posts about google accounts being closed without any recourse, I hope I'm protected of that with AWS)
Right now I only have CalDav/CardDav, my password manager and my configs being backed up, but I plan to use Syncthing to also backup other devices towards the home server, to fit inside what I already configured.
If anyone has advice on what I did/did not do/could have done better please tell me !
At work I use Windows backup to write to empty SMB-mounted drives nightly, then write those daily to another drive on an offline Fedora box.
My super critical files are on an encrypted SD card I sometimes put in my phone when cellular connection is off, and this is periodically backed up to Glacier. The phone (Galaxy) runs Dex and can be my computer when needed to work with these files.
As others mention, backup needs more than replication. You recover from a ransomware attack or other data-destruction event by using point-in-time recovery to restore good data that was backed up prior to the event. You need a sufficient retention period for older backups depending on how long it might take you to recognize a data loss event and perform recovery. A mere replica is useless since it does not retain those older copies. With retention, your worry is how to prevent the compromised machines from damaging the older time points in the backup archive.
The traditional method was offline tape backups, so the earlier time points are physically secure. They can only be destroyed if someone goes to the storage and tampers with the tapes. There is no way for the compromised system to automatically access earlier backups. You cannot automate this because that likely makes it an online archive again. A similar technique in a personal setting might be backing up to removable flash drives and physically rotating these to have offline drives. But, the inconvenience means you lose protection if you forget to perform the periodic physical rituals.
With the sort of rsync over ssh mechanism you are describing, one way to reduce the risk a little bit is to make a highly trusted and secured server and _pull_ backups from specific machines instead of _pushing_. This is under the assumption that your desktops and whatnot are more likely to be hacked and subverted. Have a keypair on the server that is authorized to connect and pull data from the more vulnerable machines. The various machines do not get a key authorized to connect to the server and manipulate storage. However, this depends on a belief that the rsync+ssh protocol is secure against a compromised peer. I'm not sure if this is really true over the long term.
A modern approach is to try to use an object store like S3 with careful setup of data retention policies and/or access policies. If you can trust the operating model, you can give an automated backup tool the permission to write new snapshots without being allowed to delete or modify older snapshots. The restic tool mentioned elsewhere has been designed with this in mind. It effectively builds a content-addressable store of file content (for deduplication) and snapshots as a description of how to compose the contents into a full backup. Building a new snapshot is adding new content objects and snapshot objects to the archive. This process does not need permission to delete or replace existing objects in the archive. Other management tools would need higher privilege to do cleanup maintenance of the archive, e.g. to delete older snapshots or garbage collect when some of the archived content is no longer used by any of the snapshots.
The new risk with these approaches like restic on s3 or some ZFS snapshot archive with deduplicative storage is that the tooling itself could fail and prevent you from reconstructing your snapshot during recovery. It is significantly more complex than a traditional file system or tape archive. But, it provides a much more convenient abstraction if you can trust it. A very risk-averse and resource rich operator might use redundant backup methods with different architectures, so that there is a backup for when their backup system fails!
Full and incremental backups of a directory tree to S3 objects, one per backup, and access to existing backups via FUSE mount. With a bit more scripting (mostly automount) and maybe shifting some cached data from RAM to the local file system it should be fairly comparable to Apple Time Machine - not designed to restore your disk as much as to be able to access its contents at different points in time.
If you're interested in it, feel free to drop me a note - my email is in my Github profile I think.
Now that you reminded me, it might be best to buy a new larger hard drive if there are any pre-Christmas sales.
For backup I use hourly & daily kopia backups that are then rcloned to an external drive and Backblaze.
Restic is darn good too! It has integration with many cloud storage providers.
Reddit also says rsync.net will accept a zfs send.
I use this to keep a few machines synced up. Including a machine that does proper daily backups.
You should be able to wake them up remotely with wake-on-lan.