Bup: Efficient file backup system based on the git packfile format
github.com
github.com
So simplified it's like: rsync -avx remote:/etc /backup/ && zfs snapshot backup@`date`
With zfSnap (https://github.com/graudeejs/zfSnap) you can tell how long incremental backups/snapshots are kept, "rsync && zfSnap -d -a 1w backup"
You can take advantage of the /backup/.zfs/snapshot directory to access all snapshots, built-in compression and possible data deduplication.
If you also have ZFS on the remote host, you can use zfs send and zfs receive to transfer the snapshot directly to the backup server, instead of using rsync for the diff.
$> ssh rsyncnet ls .zfs/snapshot
daily_2014-02-09
daily_2014-02-10
daily_2014-02-11
daily_2014-02-12
daily_2014-02-13
daily_2014-02-14
daily_2014-02-15
They allow you to customise the length of time for which these snapshots are kept, too (IIRC) (at the cost to you of the incremental extra storage)All accounts get 7 days (as the parent shows) and all 1TB+ accounts get 7 days + 4 weeks.
However, if you want a custom schedule (say, 30 days, 8 weeks, and 6 months) it only costs more if the space they use exceeds your existing account.
So, if you have a 100 GB account and you use 60 GB and your fancy snapshot schedule (of which the first 7 days is always free) only uses 35 GB ... then there is no additional cost at all.
Also, note that 1TB accounts are 15c per GB, per month, and 10TB accounts are 7c per GB, per month - all with no traffic or usage costs.
This compares very favorably with S3 and blows the mozy pro pricing out of the water[1].
https://news.ycombinator.com/item?id=6554313
(I happen to get a discount for being a prgmr.com user, but IIRC the terms are similar.)
Is there any support for that? (Of course, for a large enough hard drive, it's not much of a problem...)
http://www.synctus.com/ddar and http://github.com/basak/ddar
It's recently been made available on Homebrew, too.
That's right.
It's a simple tool that does a simple job. It's pretty much done.
> How reliable is it currently?
No known bugs.
> bup currently has no way to prune old backups
Thanks for the rdiff-backup shout-out. I'm looking for a nice way to do system backups to my NAS of large VM images without having to install Crashplan. Bup and rdiff-backup both look pretty good.This does deduplication, but can also encrypt the deduplicated blocks and store them on (say) S3.
Anything is possible with money, of course, but how is this anything other than really expensive?
For example AWS S3 would be $235/month (that's $2,820/year!) for 3TB not even including any data-out transfer charges. Sure there are others that are cheaper but only marginally so.
Is this really what people are doing? Makes the commercial services sound really cheap.
A mass restoration is expected to be rare, so it's okay for it to be a bit more expensive.
edit: 45tb for 300€/month: https://www.hetzner.de/en/hosting/produkte_rootserver/xs29
If so, that's still less than €8/TB/mo, which is better than most cloud storage providers offers. You also have some spare memory and CPU resources (so you could resell them for others as, for example, memcached instances) and a possibility to get a proper SLA, as a bonus.
1. Regularly back up "important" directories (code/, papers/, web/, etc.) to fairly safe/redundant cloud storage with incremental history. I have pretty little of this, <50gb, so it's not super-expensive.
2. Occasionally exchange bulk but less-important backups with my brother, so we're each the other's high-latency, questionable-durability "off-site backup". No incremental dumps here, just rsync. This is where my MP3 collection, DVD rips, and similar goes.
3. Photos, which are important but also bulk, go to Flickr, which is free.
4. Don't back up stuff I can re-acquire, e.g. big public datasets I've downloaded to work on, or Debian ISOs. Also, I don't back up the OS, just my data.
There do, however, seem to be some cloud services that offer big full-disk backups for a surprisingly low flat price, e.g. http://www.backblaze.com/ is $5/mo/machine.
For backups, you're dealing with a relatively consistent or predictable amount of data. Buy the appropriate dedicated server for your needs.
What I'd consider is essentially the Crashplan model: P2P / external backups locally (i.e. full LAN speed) and an off-site replica which can be cheaper and slower as long as you have a high confidence that it'll be available eventually. This way normal operations are fast but if the building burns down you're covered and presumably have higher priorities than waiting for a restore to run.
* on the webpage, there is no new release since 2009.
* has no de-duplication.
I was considering moving my 5+ years old rdiff-backup system to any of those new, promising programs:
* obnam [http://liw.fi/obnam/]
* attic [https://pythonhosted.org/Attic/]
They both do automatic de-duplication, old backup deletion and remote encryption.
I'm using obnam now.
I backup locally using duplicity, then I ship those files off to Amazon Glacier using mt-aws.
I once did a presentation about python performance optimization lessons from bup: http://lanyrd.com/2011/pycodeconf/sghxk/
And it's true, I'm not the most active maintainer anymore. The people who took over seem to be doing a pretty good job though.
I know that I would humbly submit at the least that my position has moved to believing that PyPy is a viable option for high-speed code (albeit in substantial part due to better interaction with C, nowadays).
Alternately, CrashPlan and other consumer-style services have a bad habit of using very slow, heavy, world-slowing systemwide file update scanning. :/
Having said this, a discussion of the merits and flaws of Tarsnap and similar backup services is something I'm fairly certain I've seen lengthy discussions of on similar HN posts.
(https://news.ycombinator.com/item?id=5767116 is a good source for lots of that sort of discussion)
why on earth isnt there already a perfect cross platform open source backup program? :)
I know,i know..why dont i make one myself? because we dont need a nother half done solution :-b
Killer for many commercial, overpriced services.
It may be that some languages or runtimes host consistently more reliable software than others, but I'd bet that the individual programmer, coding style and practice have more of an effect on reliability.