Zsync: Differential file downloading over HTTP using the rsync algorithm (2010)
zsync.moria.org.uk
zsync.moria.org.uk
One of the things we did to improve the efficiency of the core rolling hash based chunking logic was to add file type aware chunking. Basically, we teach the patcher about the different archive formats that our games use (say, Unreal Engine PAKs) and the algorithm ensures that chunks stop at the file boundaries within the archive, so a chunk doesn’t straddle two unrelated entries. That way when files move around in the archive we can more reliably match their chunks across versions. I think this helped improve download sizes by 5-10% in our case.
Of course for that you have to specially prepare how to store the files within the ZIP: For example the kernel image is stored, not compressed at it itself is already compressed. gzip based files (like initramfs.gz) are created with the rsyncable flag. And so one for other file types.
A really nice feature of zsync is that you can add any number of local data sources when syncing. It then uses all those to search for matching local data before retrieving it remotely. All without any explicit local state or special server side preparations. I usually add the previous and current OS install zip, the currently extracted OS files and previous interrupted transfers. zsync then magically uses all those to reduce download size as much as possible. This also means that the OS can initially reconstruct its own full install.zip file by using its extracted files and fetching the few remaining data chunks remotely.
If it swapped to libcurl as a backend, it would be much more flexible and stay up-to-date, since the idea is solid. That being said there are many spiritual successors to look into:
- zchunk
- casync
- desync
- bita
Yup. Last update of the Github repo was 2015.
zsync is effectively abandonware.
There's https://github.com/probonopd/zsync-curl. Seem pretty active. So there's hope.
(The AppImage docs suggest using zsync as a mechanism for making images updateable: https://docs.appimage.org/packaging-guide/optional/updates.h...)
Artistic License 2.0 [1],
>The Artistic License is a software license used for certain free and open-source software packages, most notably the standard implementation of the Perl programming language and most CPAN modules, which are dual-licensed under the Artistic License and the GNU General Public License (GPL).
>The Artistic license 2.0 is also notable for its excellent license compatibility with other FOSS licenses due to a relicensing clause, a property other licenses like the GPL are missing.
Hopefully saves everyone from Googling. At least that was my immediate reaction when I read the page.
There’s also a version of zsync somewhere that uses libcurl and allows sending custom headers which allows for example sending cookies to enable zsync to handle images on servers requiring custom authentication.
sendto(4, "GET /test HTTP/1.1\r\nUser-Agent: zsync/0.6.2\r\nHost: localhost:8000\r\nReferer: http://localhost:8000/test.zsync\r\nRange: bytes=1048576-1050623\r\nConnection: close\r\n\r\n", 161, 0, NULL, 0)
sendto(4, "GET /test HTTP/1.1\r\nUser-Agent: zsync/0.6.2\r\nHost: localhost:8000\r\nReferer: http://localhost:8000/test.zsync\r\nRange: bytes=0-2047\r\nConnection: close\r\n\r\n", 152, 0, NULL, 0) = 152
In both cases it transferred the only relevant 2048 byte chunk resulting from the blocksize setting of zsyncmake. Would the casync transfer be smaller than that?
Server-side maintenance/generation of .zsyncs might be transformative wherever replication is desirable, no?
1) A checksum header in the HEAD response
2) Altering the HEAD method to optionally accept the same range request headers that a GET does, and have the checksum header return the checksum for that range.
I'm mildly surprised it hasn't happened. Pretty sure you could do it with vanilla php or the nginx lua module.
Unfortunately, there's no standard as to how that's computed, so it wouldn't be useful in an rsync style setup unless you controlled both ends.
For 1)
You have SRI: https://developer.mozilla.org/en-US/docs/Web/Security/Subres... You have Etag and you have Content-MD5: https://www.w3.org/Protocols/rfc2616/rfc2616-sec14.html#sec1...
Like you already mentioned etag sucks, because you just get a guarantee X hasn't changed but no way to now how it was calculated. i.e. it is opaque.
SRI is only useful for everything outside of the original document(i.e. for <script> tags and the like.
Content-MD5 actually works.. except it's using the broken MD5 digest.. so it might work for the use case, depending on the details of the use-case. :)
2 is a no-go as far as I'm aware.
Though you could perhaps do an HTML document where you link to each "chunk" of the file using SRI.. though that would be a total bastardization of HTML and a total hack.. and a bad one at that :) Also the chunks would be decided well ahead of time that way, so not the same as a RANGE..
rclone is basically rsync with cloud native endpoints, it's not browser based but does what needs to be done very well. It does have an experimental web GUI but it's not a feature I've tried personally. https://github.com/rclone/rclone-webui-react