Chunking Attacks on File Backup Services Using Content-Defined Chunking [pdf]
daemonology.net
daemonology.net
https://www.reddit.com/r/crypto/comments/7imejm/monthly_cryp...
I've tried to take a stab at this problem, but was not sure if it worked at all:
https://gist.github.com/dchest/50d52015939a5772497815dcd33a7...
It's a modified BuzHash with the following changes:
- Substitution table is pseudorandomly permuted (NB: like Borg).
- Initial 32-bit state is derived from key.
- Window size slightly varies depending on key (by ~1/4).
- Digest is scrambled with a 32-bit block cipher.
I also proposed adding (unspecified) padding before encrypting chunks to further complicate discovering their plaintext lengths. Glad to see I was on the right track :)
Seems like that’s possible[1] to do in a fairly straightforward manner, the question is if you can do this without computing a PRF for each byte.
[1] Obviously you’re always going to leak the total data size and the approximate size of new data per each transfer.
A fresh backup will be uploading thousands of blocks. You don't want to create all the blocks before uploading, but a buffer of a hundred might be enough?
I was originally planning on having that as part of 1.0.41, but the implementation turned out to be harder than I expected.
> It seems like compression as default (or even required) is important. Without compression, Borg and Restic are susceptible to known plaintext attacks. With compression, we still have theoretically sound (and harder) chosen-plaintext attacks but no known-plaintext attacks. Sadly, compression can also leak information for post-parameter extraction attacks, as shown in Example 4.3.
EDIT: or maybe this keyed rolling hash https://crypto.stackexchange.com/questions/16082/cryptograph...