Time-Machine-style backup with rsync
github.com
github.com
Without hard linked directories, a full --link-dest backup of a decent sized disk, with zero file changes from the previous pass, can easily consume 100MB (and take 45 minutes to perform).
This disk consumption might seem insignificant, today, but that's 2.4GB per day if you run a standard Time Machine equivalent backup schedule. Of course you might not choose to do that, because the previous hour's backup would only finish 15 mins before the next one started, which is insane.
These numbers are from direct experience on a 2TB, approximately 60% utilized source drive.
That said, I use rsync, not Time Machine, for my OSX backups. You'll want a few additional switches for HFS+, and if your target drive is HFS+ also, make sure you turn OFF "ignore ownership on this volume" in the Finder... but the script posted here has the right general idea. Somewhere on my project list is adding directory hardlinking to rsync.
With hard-links the underlying data is not deleted until all hard-links are deleted, so you can delete any individual backup directory without losing data in any other backup directory.
A soft-link is like a pointer in C whereas a hard-link is like a C++ shared_ptr, ie. reference counted.
Another way, which is worse, is block-level deduplication. All of the above filesystems support it, as does NTFS.
I wish Apple would adopt HAMMER for Mac OS. It is BSD-licensed and more suitable for a memory-constrained environment than ZFS.
Time Machine, as can be seen here http://www.apple.com/support/timemachine/ time machine is tightly integrated in the os and provides a self-defining interface and user experience.
This github page is for a wrapper shell script around rsync, which is not like time machine.
This has been going on for a while now, see Timevault (https://wiki.ubuntu.com/TimeVault ), back in time (http://backintime.le-web.org/ ) or flyback (http://www.flyback-project.org/ ).
Please telling us your backup solution is like time machine when it lacks the kind of UX time machine offers, thanks !
http://www.youtube.com/watch?v=RDPzVdohrck#t=1m13s
Oh, and integration with the OS X recovery partition / OS reinstallation mechanism that allows you to point to a Time Machine backup as the recovery point for your re-installation.
TM has had a few problems, but by and large it is one of the quiet successes in OS X, and probably my favorite feature if the OS. Why Microsoft hasn't put something like it in Windows is baffling to me.
Be sure to test this to make sure it restores, but in my case it works flawlessly.
I've become obsessed with the "12factor" app approach everywhere in my digital life, so that if a device ever disappeared, was stolen, or died, I could get a replacement fully operational without any problems. Like an app-server dying, just launch a new one and it will bootstrap itself.
I'm kinda crazy with my new "homelab" and started it off with an old 2U Poweredge I got on eBay for about $200. It's cheaper than a Synology/Drobo, has room for 6 drives, and the dual quad-core processors + 16GB of RAM is pretty cool too. It's running FreeNAS right now in a VM with a 3TB ZFS pool. I've created an AFP share that appears to my mac as Time Machine and over my gigabit-network it does a pretty fast backup. I really like the FreeNAS software. It's open-source, runs on FreeBSD, and the UI/admin tool is built in Django.
You should also check out rsnapshot: http://www.rsnapshot.org/, it does a great job and has many of the features that this script does.
I've used dirvish for several years (nearly a decade), both locally and remotely without issue. No configuration is necessary on the client, only on the backup server (alternatively you can flip it and put all of your configuration on the client). Uses rsync, hard links, and can keep snapshots at various intervals.
I started too with hard links but nowadays I prefer to format my backup disk with BTRFS and use btrfs snapshots instead, though I still support hard links.
I prefer btrfs snapshots due to their support for COW, so if I decide to play with a backup I won't mess all the other versions of this backup. With hard links you should never write to your existing backups.
Lately I added a helper script to mount remote filesystems and lock mysql databases but it could be easier to use.
Most of my code is checks to make sure I won't write somewhere I shouldn't to.
You should, however, consider bup (https://github.com/bup/bup) - it takes less than a minute to figure out nothing is done, it deduplicates parts of files, (that is, if you have a 20GB virtual machine image, and you've changed one byte in the middle of it, then the next snapshot is going to take ~10KB, not 20GB). The older release don't keep ownership/modification time, but there's a new version pending release soon that does.
It also works well remotely (through ssh), can do an integrity check (bup fsck), redundancy (using par2; important after deduplication). And it has a fuse frontend that makes it all accessible as a file system, as well as an ftp frontend.
bup is teh awesome.
Also, if you're doing multiple backups of the same data to a filesystem over time, it's worth doing 'cp -al' from the previous backup to the current backup destination, then rsync over the top of that - that way multiple copies of files which haven't changed don't take up any extra space.
Isn't that already taken care of by its use of rsync's --link-dst option?
Doesn't vim intentionally do this if it edits files that are hard linked? emacs breaks the hard link. I'm not saying either behavior is best, but in this case one might be surprising!
What I don't like about that solution is it can't dedupe (ZFS dedupe just ultimately doesn't work very well). Hence my interest (and if anyone checks my post history, my constant spruiking of) bup - which does efficient dedupe and output of git pack-files. Stick that on a ZFS volume with snapshots, and you've got block-level checksummed, versioned and deduplicated backups.
What it's all missing of course, is a pleasant interface to use it with (one which doesn't fallback to the thing I see way too often in a lot of these scripts "don't worry, we're just going to stat your entire filesystem every 20 minutes).