How does rsync work?
michael.stapelberg.ch
michael.stapelberg.ch
On the question of what happens if a file's contents change after the initial checksum, the man page for rsync[0] has an interesting explanation of the *--checksum* option:
> This changes the way rsync checks if the files have been changed and are in need of a transfer. Without this option, rsync uses a "quick check" that (by default) checks if each file's size and time of last modification match between the sender and receiver. This option changes this to compare a 128-bit checksum for each file that has a matching size. Generating the checksums means that both sides will expend a lot of disk I/O reading all the data in the files in the transfer (and this is prior to any reading that will be done to transfer changed files), so this can slow things down significantly.
> The sending side generates its checksums while it is doing the file-system scan that builds the list of the available files. The receiver generates its checksums when it is scanning for changed files, and will checksum any file that has the same size as the corresponding sender's file: files with either a changed size or a changed checksum are selected for transfer.
> Note that rsync always verifies that each transferred file was correctly reconstructed on the receiving side by checking a whole-file checksum that is generated as the file is transferred, but that automatic after-the-transfer verification has nothing to do with this option's before-the-transfer "Does this file need to be updated?" check. For protocol 30 and beyond (first supported in 3.0.0), the checksum used is MD5. For older protocols, the checksum used is MD4.
Newer versions (≥3.2?) support xxHash and xxHash3:
* https://github.com/WayneD/rsync/blob/master/checksum.c
* https://github.com/Cyan4973/xxHash
* https://news.ycombinator.com/item?id=19402602 (2019 XXH3 discussion)
Manpage for rsync 3.2.4, from Debian testing: https://manpages.debian.org/testing/rsync/rsync.1.en.html
* Underlying disk device corruption - but modern disks do internal error checking, and should emit an IO error.
* Corruption in RAM/software bug in the kernel IO subsystem. Should be detected by filesystem checksumming.
* User has accidentally modified file and set mtime back. fixes this case.
* User has maliciously modified file and set mtime back. Since it's MD5 (broken), the malicious user can make the checksum match too. checksumming doesn't help.
Given that, I see no users who really benefit from checksumming. It isn't sufficient for anyone with really high data integrity requirements, while also being overkill for typical usecases.
No, md5 is not broken like that (yet). There is no nkown second-preimage attack against md5; the practical collision vulns only affect cases where an attack controls the file content both before and after update.
Also, filesystem checksumming does not guard against ram/kernel-bugs. On top of that file system checksumming is very rare.
For anyone curious about how to find such problems without changing the files, I used "--checksum --dry-run --itemize-changes"
Which is an actual case that has occured for myself.
It depends. I recently built a new zfs pool server and needed to transfer a few TB of data from the old pool to the new pool, but I built the new pool with a larger record size. If I’d used zfs send the files would have retained their existing record size. So rsync it was.
IIRC neither ZFS nor btrfs use cryptographic hashes for checksumming by default.
* https://openzfs.github.io/openzfs-docs/Basic%20Concepts/Chec...
* https://people.freebsd.org/~asomers/fletcher.pdf
* https://en.wikipedia.org/wiki/Fletcher%27s_checksum
Strangely enough SHA-512 is actually (50%) faster than -256:
> ZFS actually uses a special version of SHA512 called SHA512t256, it uses a different initial value, and truncates the results to 256 bits (that is all the room there is in the block pointer). The advantage is only that it is faster on 64 bit CPUs.
It's also hard/impossible to restore individual files out of a ZFS send stream without restoring the whole thing so I've reverted to using tarballs of ZFS snapshots for backups instead of ZFS send. Again, it was never really meant for this so it was my mistake trying to use it that way.
linux.die.net is horribly outdated. This particular page is from 2009.
Up-to-date docs are here:
If you want to see the man page of the version in Debian, that would be https://manpages.debian.org/testing/rsync/rsync.1.en.html
Disclaimer: I wrote the software behind manpages.debian.org :)
[0]: https://blog.liw.fi/posts/rsync-in-python/
https://web.archive.org/web/20150321212547/http://code.liw.f...
* https://rsync.samba.org/tech_report/
* https://www.andrew.cmu.edu/course/15-749/READINGS/required/c...
I checked the three individually but they showed no corruption or changes either side. How can this happen?
Edit: the hard drive copy is previously rsynced from this copy & both copies are mirrored with google cloud bucket.
The 3 files which showed different have the same MD5 checksum
Maybe it's a trivial thing, eg. your 3 files got resynchronized right between running rsync and running diff. So you should have retried rsync after the diff.
Or you obtained these PDFs from a source that purports to demonstrate MD5 collisions. Or someone is attacking you by replacing your files. Or, more likely, this is user error, and you are not reporting to us what's happening exactly.
You can always do a diff on a hex dump of the PDF content and see with your own eyes what part of the PDF is actually different. It's not that hard to interpret the format and know which PDF structure changed. You can run "qpdf --qfd input.pdf" on both versions and this uncompress all structures to make the internal content human readable (besides images).
>Or, more likely, this is user error, and you are not reporting to us what's happening exactly.
Here's the sequence:
1. rsync -av src/ dest/
2. diff -r src/ dest/
[Shows 3 pairs of differences]
3.
md5 file1a & file 1b; compare
md5 file2a & file 2b; compare
md5 file3a & file 3b; compare
[All three pairs match MD5]4. run rsync - no difference still
5.compare difference by diff - shows gibberish. All three copies open by pdf viewers. The qpdf option doesn't make sense because all the 3 happen to be advanced math textbooks, and the plaintext is impossible to read.
In fact I just checked now, and the same error pattern persists. Its not something I am terribly concerned since I have 3 copies of my data - but this is a pattern which showed up for the first time. I do this exercise regularly (once per month)
It's possible to have a scenario where your synchronization process (involving Google Cloud?) somehow changed the modification time of your files so that both copies in src/ and dst/ have the same timestamp, in that case rsync will not notice they are different (if they also have the same size). Like the other reply said you have to use rsync --checksum or -c to force rsync to compare the content of files.
1. rsync -acv src/ dest/
It will be much slower - using lots of cpu and disk access - but more thorough.
Just last week, Dropbox unilaterally decided I didn't want local copies of the shared files on my laptop, which made for some awkwardness inside a secured facility with no Internet access.