Casync – A Content-Addressable Data Synchronization Tool
github.com
github.com
[0] https://github.com/SaveTheRbtz/zstd-seekable-format-go
[1] https://github.com/facebook/zstd/blob/dev/contrib/seekable_f...
[2] e.g FastCDC https://www.usenix.org/system/files/conference/atc16/atc16-p...
This change achieved a 10X speedup on normal operations compared to the old code that used SCCS files.
The compressed format stored data blocks in arbitrary order and then the index at the end of the file gave the data layout. Then allows write to the file without rewriting. BitKeeper itself only needed to append to the end of the file, but the format could support inserting data in the middle by only appending to the physical file.
It also had a data redundancy CRC layer that could detect data corruption and recover data from some types of corruption.
https://www.bitkeeper.org/ https://github.com/bitkeeper-scm/bitkeeper/blob/master/src/l... https://github.com/bitkeeper-scm/bitkeeper/blob/994fb651a404...
Another thought is if it is possible and how to coordinate te re-using dictionaries.
If you need a mature compression implementation for bazel I would recommend using recent bazel versions w/ gRPC-based bazel-remote: https://github.com/buchgr/bazel-remote
bazel nowadays supports end-to-end compression w/ `--experimental_remote_cache_compression`: https://github.com/bazelbuild/bazel/pull/14041
But looking at the code I'm having strong "nope" feelings. First, because of lines like "q += m, n -= m;". Second, because of int/enum/semantic abuse: `compression_type` may be _CA_COMPRESSION_TYPE_INVALID which I hope is -1, `>= 0` as a known compression type, or `-EAGAIN` as an error. (from https://github.com/systemd/casync/blob/99559cd1d8cea69b30022... ) I'd bet that just throwing afl at the decompressor will find issues :( (source - I had this feeling about systemd-resolved and threw afl at it, found issues)
I do like the idea though.
When I have big 100MB binary files that i want to version, the changes are small (1MB in one place, and a few more KB in others). I also have a few multi GB SQLite databases I would like to version where this would help. (Changes are less than 5% of the data, but index pages throughout the file mean rsync-style partitions still transfer 50% or so of the file. To actually achieve storage efficiency, I store textual dumps, and they also fit better with borg/restic/casync than git or git-lfs
Would have been extra awesome if desync would be able to use a git repo as storage.
For text files with really small changes, this is comparable to a "diff" size, and is better than what a rolling hash would achieve. However, for text files with larger changes, and of course for binary files - a rolling hash is much more effective. Additionally, a rolling hash would easily reuse cross-file similarity, where as the current delta-finding code is likely to only find redundancy in the history of the same file (or similarly named in the same directory).
> Is this a systemd project? — casync is hosted under the github systemd umbrella, and the projects share the same coding style. However, the code-bases are distinct and without interdependencies, and casync works fine both on systemd systems and systems without it.
Didn't they say that about udev?