Checksumming drives a lot of support calls that would otherwise not happen. As long as the errors are in media files and not meta-data, most consumers are going to be oblivious to bit rot in their downloaded movies and photos.
Enabling checksumming is going to reveal a lot of errors that would otherwise be silently ignored (spoken from experience running a ZFS media server)
ReFS already does this when you're running it on non-redundant storage - integrity streams are only enabled on mirrored/parity configurations.
It would be unbelievable if they really didn't have checksumming. Could it just be that they haven't documented it yet? That seems weird and unlikely, but... not was weird and unlikely as APFS not having checksumming.
Rather than having to RAID-mirror every block on my disk, I'd like to be able to pick just some files and say "please store those ones slightly more redundantly, such that they're protected from bit-level disk errors—they're important."
zfs set checksum=off rpoolWhen a sector goes bad on a hard drive, the firmware will mess about retrying and altering the analog amplifiers to try to get the signal back. If it gets the data back, it might "recover" by moving the data to spare sectors. Without a CRC, the firmware would have no way of differentiating between a sector read correctly and one that is read erroneously.
If block 1027 is errantly overwritten due to a bug in the filesystem, the block driver, the DMA subsystem, or the device firmware, the FS will know as soon as it goes to fetch block 1027 that something went wrong. This is true even if the block is still internally consistent; perhaps the write was misdirected, another sector was incorrectly read, or a failure in the flash translation layer caused a newer write to be lost.
If there's any redundancy to the system, whether storing the metadata in multiple places on the disk, or on different disks in an array, the FS can then check all the others and, crucially, detect which version is correct.
I'll gladly spend 1% or whatever space overhead to notice it.
FS level checksums also cover bitflips on your bus during the write phase (the drive will crc the already bad data) or during the read phase (you receive different data than your drive sent).
If I see one bad sector develop on a disk for no obvious reason, it's getting replaced ASAP because it's a sign that more will follow very soon.
Btw: So far the Backblaze reliability numbers check out (<1% annual failure rate)