> If you ran into a problem like this on ZFS, for example, you'd have very high confidence about whether the disk was at fault
Would we, though? I'll admit to not being that familiar with ZFS's internals, but I'd be a bit surprised if its checksums can detect lost writes. More generally, I'm not entirely sure how practical it would be to add verification at all layers of the stack, as you seem to be suggesting.
We'd certainly be open to considering ZFS in future if it can help track down this sort of problem.