Sadly, the only thing I've found so far that works at all at those scales is Bacula, and that is file-based -- i.e., if you have a gigabyte file that changes by one byte, it backs up the whole gigabyte again. Not ideal.
Sadly, the only thing I've found so far that works at all at those scales is Bacula, and that is file-based -- i.e., if you have a gigabyte file that changes by one byte, it backs up the whole gigabyte again. Not ideal.
Restic falls over as soon as you cross 2TB sizes, and prune operations are painfully slow. Backing up 8TB took 4 weeks with Restic, and i gave up waiting for prune to finish. It was crossing the 24 hour mark, making it unsuitable for daily backups.
Would it possible to split those bigger files before doing backup (log files with log splitters, sql databases with incremental specialised backup tools like e.g. xtrabackup), or are these e.g. image files with different versions?
I really like the concept of byte level deduplication, but have often thought the price you pay for the space reduction might not be worth it considering today's network speed as well as storage sizes and prices.
Would be interesting to hear your experience concerning this!
If the backup service is using file timestamps as the only key to refresh then it would have to store the last modification date and a good hash of each block, and when the date of the file being mirrored is updated scan each block to see if there is a change there and update the stored hash and timestamps accordingly. This would need to be orchestrated to reduce the risk of temporary corruption if there are several updates and the rechecking process coincides with a backup sweep (i.e. make sure you don't present updated dates for any block until you can present them for all needed).
For restoration, you either manually concatenate the parts or have an overlay filesystem that operates in reverse: showing the smaller block files as a single large unit.
You'd have to very thoroughly test the overlays and their interaction with the backup service before risking it on important data, so it might not be something you would genuinely consider...
But I'm not very happy about it, because it's insanely overcomplicated for no really good reason, and because of it being file-based.
https://www.backblaze.com/blog/lto-versus-cloud-storage/
If you're interested, I'm doing experiments with another site with 500T to backup, where I added sampling and sharding to HashBackup (I'm the author).
Sampling allows you to do faster simulated backups to determine the best backup parameters to use. In his case, we determine that a very large block size - 64M - was the best way to backup his data.
Sharding automatically partitions the filesystem so multiple backup can run simultaneously to get backup speed in the 250-400 MB/s range.
It's more at the proof of concept stage, but having another large site to work with would be fantastic! A couple of the larger sites using HashBackup are EURAC (European Research Center) and HMDC (Harvard MIT Data Center)