Zsync: Differential file downloading over HTTP using the rsync algorithm (2010)
zsync.moria.org.uk
zsync.moria.org.uk
Its features make it very handy for a number of file transfer/sync tasks, over and above its chief SCM role.
That said, I'm not sure I know of any other major users of it -- most people just use a .torrent (which similarly has checksums of each piece so you know which pieces need to be downloaded).
On the other hand, HTTP has the advantage that it works through corporate proxies and that is usually not blocked by over-cautious admins.
- HTTP(s) based trackers (although UDP is more common these days) - HTTP webseeds/mirrors (BEP-17) - (if you count it) webtorrent uses Websocket trackers and can support HTTP webseeds (although really of course the P2P is WebRTC)
Zsync is awesome as we can specify any number of existing files already available on the device (with the -i command line option) and zsync will try to make use of them to minimize downloads. We really use this feature to our advantage: zsync by default will keep the previous version of a file if it's going to overwrite it. So we have two versions of install.zip on a device. When switching between OS releases (stable / testing...) we can switch back and forth with zero additional downloads as both versions are available and zsync makes use of that. Similarly after a user installs our OS, we just have the unpacked artefacts (kernel, etc.) on the SD. We can quickly recreate an initial version of the install.zip file on the device by seeding the download with those files. It's usually just 500k to construct an initial install.zip file we then later use to minimize all future updates.
https://sourceforge.net/projects/antix-linux/files/Final/ant...
Can you achieve differential downloads with bittorrent?
If you are looking for a maintained system for online systems that provide software updates I would look into https://github.com/itchio/wharf-spec.
Wharf is used for Itchio to sync folder structures differentially / incrementally. It uses the latest compression algorithms. It has a reference server.
Alternatively, a port of zsync is https://github.com/salesforce/zsync4j written in Java.
I had trouble compiling zsync for windows.
https://itch.io/docs/wharf/algorithms/diff.html and https://itch.io/docs/wharf/algorithms/apply.html describe patching
The most important thing is that Itchio runs a business on the usability of this system.
How does zsync diff against the local file without downloading the contents from the server?
It works by expecting a .zsync metadata file which gets downloaded first and is used to guide the differential download.
"Rsync over HTTP — zsync provides transfers that are nearly as efficient as rsync -z or cvsup, without the need to run a special server application. All that is needed is an HTTP/1.1-compliant web server. So it works through firewalls and on shared hosting accounts, and gives less security worries."
The other child comment also touches on something that I seem to have skipped over, the .zsync metadata file which contains pre-calculated rsync hashes. Using this means software isn't needed, but during upload the file needs to be processed to produce this metadata.
> in fact, the .zsync can be generated and offered by a third party, while still leaving most of the download to the original distribution site.
So you'd zsync http://third-party/file.zsync and this file would contain the url for the main file.
> What happens if the .zsync is out of sync with the resource?
I'm not sure, I've not managed to get it working on a simple file yet.
It's just a pair of files. The big thing you're trying to transfer and the zsync file that details the content of that first file, to guide the downloader.
You don't need a plugin, you don't need to change your httpd.conf, ...
It works as follows:
- Let say you already have file that is older version, or perhaps corrupted, you use rdiff to generate its signature
- you go then to the place which contains proper file and use the signature file to generate a patch file
- then you use the patch file to fix your local file
There is an elegant simplicity to rdiff, IMHO.
rdiff-backup backs up one directory to another, possibly over a network. The target directory ends up a copy of the source directory, but extra reverse diffs are stored in a special subdirectory of that target directory, so you can still recover files lost some time ago. The idea is to combine the best features of a mirror and an incremental backup. rdiff-backup also preserves subdirectories, hard links, dev files, permissions, uid/gid ownership (if it is running as root), and modification times. Finally, rdiff-backup can operate in a bandwidth efficient manner over a pipe, like rsync. Thus you can use rdiff-backup and ssh to securely back a hard drive up to a remote location, and only the differences will be transmitted.
Unfortunately, it seems not to be developed any longer, and it has a few things that would need ironing out:
* You can't pause a backup and continue later. * Some operations (notably recovery after an aborted backup run) is excruciatingly slow. It takes tens of hours for me with a backup of 40 GB or so (on a low-powered computer as server, though). I think rdiff-backup-fs is resource hungry as well, which is perhaps partly understandable, since it has to go through a series of reverse diffs to present old versions of a file. * I tried it on Windows once, and it could apparently not handle paths longer than a few hundred characters (due to using that older Windows API, whatever it's called). * You can't delete intermediate backups, only the oldest one.
zsync also does all the fetching stuff directly.
> Not sure how SCP would help here.
I mean that my train of though was first "scp, it seems, can do everything cp can, and more, so let's drop cp and use scp instead", then "but hey, rsync, it seems, can do everything that scp can, and more, so let's drop scp too and use rsync instead".