Given I worry about this sort of thing for a living and am a partner in the firm: I think in terms of backup, DR, BC, availability and more. I have access to rather a lot of gear but the same approach will work for anyone willing to sit down and have a think and perhaps spend a few quid or at least think laterally.
For starters you need to consider what could happen to your systems and your data. Scribble a few scenarios down and think about "what would happen if ...". Then decide what is an acceptable outage or loss for each scenario. For example:
* You delete a file - can you recover it - how long
* You delete a file four months ago - can ...
* You drop your laptop - can you use another device to function
* Your partner deletes their entire accounts (my wife did this tonight - 5 sec outage)
* House burns down whilst on holiday
You get the idea - there is rather more to backups than simply "backups". Now look at appropriate technologies and strategies. eg for wifey, I used the recycle bin (KDE in this case) and bit my tongue when told I must have done it. I have put all her files into our family Nextcloud instance that I run at home. NC/Owncloud also have a salvage bin thing and the server VM I have is also backed up and off sited (to my office) with 35 days online restore points and a GFS scheme - all with Veeam. I have access to rather a lot more stuff as well and that is only part of my data availability plan but the point remains: I've considered the whole thing.
So to answer your question, I use a lot of different technologies and strategies. I use replication via NextCloud to make my data highly available. I use waste/recycle bins for quick "restores". I use Veeam for back in time restores of centrally held managed file stores. I off site via VPN links to another location.
If your question was simply to find out what people use then that's me done. However if you would like some ideas that are rather more realistic for a generic home user that will cover all bases for a reasonable outlay in time, effort and a few quid (but not much) then I am all ears.
Someday I'm going to get a second offsite system to do ZFS backups to, but so far the above has served well. Then again I've been lucky enough to never have a hard drive fail, so the fact that I can lose 2 without losing data is pretty good. I'm vulnerable to fire and theft, but the most likely data loss scenarios are covered.
I've got incremental snapshots sent to a box in a family member's house (and vice versa).
"rdiff-backup backs up one directory to another, possibly over a network. The target directory ends up a copy of the source directory, but extra reverse diffs are stored in a special subdirectory of that target directory, so you can still recover files lost some time ago. The idea is to combine the best features of a mirror and an incremental backup. rdiff-backup also preserves subdirectories, hard links, dev files, permissions, uid/gid ownership (if it is running as root), modification times, acls, eas, resource forks, etc. Finally, rdiff-backup can operate in a bandwidth efficient manner over a pipe, like rsync. Thus you can use rdiff-backup and ssh to securely back a hard drive up to a remote location, and only the differences will be transmitted."
- daily snapshots for 1 week
- the first snapshot of every week for 4 weeks
- the first snapshot of every month for 2 years
- the first snapshot of every year for 5 years
Since I use entirely SSD storage I also have a script that mails me a usage report on those snapshots, and I manually prune ones that accidentally captured something huge. (Like a large coredump, download, etc. I do incremental sends, so I can never remove the most recent snapshot.)
Since snapshots are not backups I use `btrfs send/receive` to replicate the daily snapshots to a different btrfs filesystem on spinning rust, w/ the same retention policy. I do an `rsync` of the latest monthlies (once a month) to a rotating set of drives to cover the "datacenter burned down" scenario.
My restore process is very manual but it is essentially: `btrfs send` the desired subvolume(s) to a clean filesystem, re-snapshot them as read/write to enable writes again, and then install a bootloader, update /etc/fstab to use the new subvolid, etc.
---
Some advantages to this setup:
* incremental sends are super fast
* the data is protected against bitrot
* both the live array & backup array can tolerate one disk failure respectively
Some disadvantages:
* no parity "RAID" (yet)
* defrag on btrfs unshares extents and thus in conjunction with snapshots this balloons the storage required.
* as with any CoW/snapshotting filesystem: figuring out disk usage becomes a non-trivial problem
The important stuff (projects, dotfiles) I keep on Tarsnap. I also rsync my entire home directory to an external drive every other week or so.
Similar for servers but I do back up /etc as well.
Now keeping a few short copies is also fine provided you don't make mistakes. Have you ever wanted to recover from a cock up you did six months ago or three years ago?
You do offsite (Tarsnap), so you have covered off local failures - cool.
Everyone's needs are different and the value they place on their data is different but I would respectfully suggest that you think really hard about how important some bits of your data are and protect them appropriately.
Fuck ups are hard to recover from 8)
This is what I do as well, just not quite as often. Sometimes I wonder if I should switch to something like rdiff-backup to get snapshots, but that would only really be useful if I accidentally deleted a file and didn't notice for a while, for instance, which in practice is not a serious problem.
If I had more time and inclination, I might set up a small 2-drive RAIDed NAS box, and do automated regular backups to that. But for my laptop PC, just doing regular syncs to an external HD seems to be fine for now.
Admittedly I'm not sure that you can consider one of them that went from i386 to amd64 as the same install simply because /var/lib/portage/world and /etc/portage/ (plus a few other bits) are the same. I still use it and it's on its fifth or sixth incarnation as my current laptop.
Trigger's broom?
(OK https://en.wikipedia.org/wiki/List_of_Ship_of_Theseus_exampl...)
2 HD failures in the last 4 years. Not rare for me :-(
Only real things missing is encryption support (working on that), and backing up KVM virtual machines from the host (working on that too).
#helps to see the fstab first
UUID=<rootfsuuid> / btrfs subvol=root 0 0
UUID=<espuuid> /boot/efi vfat umask=0077,shortname=winnt,x-systemd.automount,noauto 0 0
cd /boot
tar -acf boot-efi.tar efi/
mount <rootfsdev> /mnt
cd /mnt
btrfs sub snap -r root root.20170707
btrfs sub snap -r home home.20170707
btrfs send -p root.20170706 root.20170707 | btrfs receive /run/media/c/backup/
btrfs send -p home.20170706 home.20170707 | btrfs receive /run/media/c/backup/
cd
umount /mnt
So basically make ro snapshots of current root and home, and since they're separate subvolumes they can be done on separate schedules. And then send the incremental changes to the backup volume. While only incremental is sent, the receive side has each prior backup to the new subvolume points to all of those extents and is just updated with this backups changes. Meaning I do not have to restore the increments, I just restore the most recent subvolume on the backup. I only have to keep one dated read-only snapshot on each volume, there is no "initial" backup because each subvolume is complete.Anyway, restores are easy and fast. I can also optionally just send/receive home and do a clean install of the OS.
Related, I've been meaning to look into this project in more detail which leverages btrfs snapshots and send/receive. https://github.com/digint/btrbk
We back up photos from our iOS devices to this server using an app called PhotoSync. I also have an instance of the Google Photos Desktop Uploader running in a docker container using x11vnc / wine to mirror the photos to Google Photos (c'mon Google, why isn't there an official Linux client???). I'm really paranoid about losing family photos. I even update an offsite backup every few weeks using a portable HDD I keep at the office.
No you aren't paranoid. Sensible. However, you've only just started. Try and do a restore every now and then from the HDDs.
There is no such thing as paranoia when it comes to protecting your data.
Not only do I do that, but I also occasionally pull down some files from Crashplan just to make sure everything is backing up and working as expected.
Don't forget that families/friends can also effectively offsite each other's backups without invoking third parties. That does assume connectivity, storage etc. Also it does need managing 8)
Plain old btrfs snapshot + rsync to local usb drive and offsite host for /etc, /var, /root
Duply for servers, keeping backups on S3: http://duply.net/
Cron does daily DB dumps so Duply stores everything needed to restore servers.
- Crashplan[1] for cloud backups (also nightly; crashplan can backup continuously but I don't do that)
Pretty happy with it, though dirvish takes a little bit of manual setup. Never had to resort to the cloud backups yet.
Backup takes ~10min for searching 1TB of disk space. The daily diff is typically 6..15 GB, mostly due to braindead mail storage format...
I want to keep it simple but still have full history and diff backup: no dedicated backup tool, but rsync + btrfs. A file-by-file copy is easy to check and access (and the history also looks that way).
If the source had btrfs, I would use btrfs send/receive to speed it up and make it atomic.
I have two such backup disks in different places. One uses an automatic backup trigger during my lunch break, the other is triggered manually (and thus not often enough).
The sources are diverse (servers, laptops, ...). The most valued one uses 2 x 1 TB SSDs in RAID1 for robustness.
All disks are fully encrypted.
btrfs subvolume snapshot / send are not atomic, for various definitions of atomic.
Unlike zfs, subvolume snapshots are not atomic recursively. That is, if you have subvol/subsubvol, there's no way to take an atomic snapshot of both. At least this one is obvious, since there's no command for taking recursive snapshots, so it tips you off that this is the case. Not having an easy way to take recursive snapshots, atomic or not, is a different pain point...
What's more insidious is that after taking the snapshot, you must sync(1)[0] before sending said snapshot, otherwise the stream would be incomplete! I'm invoking Cunningham's Law here and saying for the record that this is fucking retarded. I have lost files due to this...design choice.
Moreover, though is is probably a fixed bug, I used to have issues where subvolumes get wedged when I run multiple subvolume snapshot / send in quick succession. I'd get a random unreadable file, and it's not corruption (btrfs scrub doesn't flag it). Usually re-mounting the subvolume will fix it, and at worst re-mounting the while filesystem would fix it so far. I haven't had it happen for a while, but it's either due to my workaround -- good old sleep(60) mutex -- or because I'm running a newer kernel.
I can't wait until xfs reflink support is more mature: that'll get me 90% what I use btrfs for.
tl;dr: btrfs: here be dragons!
[0] https://btrfs.wiki.kernel.org/index.php/Incremental_Backup
It has an ugly looking interface, but the core of the product is super reliable.
Cron runs it, on @reboot schedule. If the backup is successful, some (but not all) old backups are deleted. I delete some oldest preserved backups manually, if disk space runs low.
Its main advantages with respect to other approaches, at least for my use cases, are:
- metadata is stored in a single append-only file: no extra software (DB etc.) is needed;
- partial backups can be performed to separate storages. In fact, source and backup directories are not conceptually different, so a duplicate of a directory counts as a backup.
I run a weekly script to rsync one HD to another. The backup HD is exactly the same size and partitioned identically. I had a HD crash some years ago and it was fairly trivial to swap out the drives (probably needed to make some changes to the MBR). Unfortunately, I had an HD crash some months ago and it was not as easy this time round. Apparently my rsync would fail in the middle and so a lot of files were stale. Unbootable. Fortunately, all the critical data was copied.
I should have a smarter backup script that will alert me on failure to rsync.
My customers and employers have all had these rube goldberg enterprisey backup systems, usually Symantec or Veritas talking to HP MSAs.
Delete one of them and then get it restored by your supplier - I assume you've done that already. I did.
I have a systemd timer to run (incremental) backups every 3 hours, and I plan on setting up a mechanism to automatically verify all of my data that has been uploaded.
At work we do regular TAR backups to external drives and SSH-rsync data to our sister office via VPN nightly. Backups are good for system restore then rsync back from remote to get to most recent.
Sidenote: I was using Time Machine on MacOS but since I upgraded to 10.13, APFS disks are mandatorily excluded by the OS (apparently as a workaround to to some bug), so restic it is too.
[0]: (warning: jwz) https://www.jwz.org/doc/backups.html
`rsync -av --files-from=".backup_directories"` on a daily cron job.
iMac and work hackintosh rsynced to local and remote backup machine daily. Windows machine (where all my music and pictures live) both rsynced to local and remote backup machine, and Cobian'd to a second drive daily. I also will run the same Cobian backup to a cold external drive every month or so.
Deathly afraid of data loss.
My work files are kept in git repos which I push to the same servers.
I use Ansible to configure my machine, so that I don't have to backup system files, just the playbooks.
2. Clone said drive to an external drive. Detach and lock it in a water/fireproof box when not in use.
3. Swap external drive with another that is stored off site every week or two.
4. Swap with yet another off site external drive less often (a few months).
- A home NAS (4x5T + 7x4T = 48T, btrfs raid1)
- A 2U server sitting in a local datacenter (8x8T + 2x4T = 72T, btrfs raid1)
- An unlimited Google Drive plan
I run periodic snapshots on both servers and use simple rsync to sync from the home NAS to the colo. Irreplaceable personal stuff is sent from the colo to Google Drive using rclone.
Synology NAS backed up daily to Backblaze B2
Dead simple to set up and maintain, and in the event that I need to restore a file or files, it's relatively fast.
Each device gets its own ZFS filesystem and is snapshotted after rsync.
FolderSync on Android does this automatically when I'm on home wifi. AcroSync for Windows. Both FolderSync and AcroSync are worth the small purchase price. Cronjobs for nix machines. iPad syncs to Mac which has a cronjob.
Stuff I really don't want to lose (photos, music, other art) are on multiple machines + cloud.
My (two) servers: dump of db, rsync of dumps and files to another server.
It's ok only because I've got little data.
Then the usual btrfs send/receive tricks.
- 6 month archival image of NAS to external HDD, rotated every 2 years
- 3 month differential rsync to nearline storage, kept for 5 years
Symlink ~/.private to Dropbox/private
Per file encryption, and I do not care if Dropbox will get hacked again
With rsnapshot, I have hourly, daily, weekly and monthly backups that use hard links for the saving disk space. These backup dirs can be mounted read-only.
With afio, files above some defined size are compressed and then added to the archive, so that if some compression goes wrong, only that file may be lost, the archive is not corrupted. Can have incremental backups.
From the afio webpage: Afio makes cpio-format archives. It deals somewhat gracefully with input data corruption, supports multi-volume archives during interactive operation, and can make compressed archives that are much safer than compressed tar or cpio archives. Afio is best used as an `archive engine' in a backup script.
[1] http://rsnapshot.org/ [2] http://members.chello.nl/k.holtman/afio.html
I also keep current copies of most configuration data for all systems, mainly by backing up their /etc directories. This is also done for network equipment and remote network configuration data (zone files, etc).