> So the read fails. And when that happens, you are one unhappy camper. The message "we can't read this RAID volume" travels up the chain of command until an error message is presented on the screen. 12 TB of your carefully protected - you thought! - data is gone. Oh, you didn't back it up to tape? Bummer!
So, at this point, we've got two hard drives containing millions of sectors for which one or two are bad, and the software claims that the whole thing is broken, and we have to find a new array?
As far as I know, RAID 5 is block level - the block level failure should only destroy one block. All of the others (apart from the other dead ones) are fine. This sort of thing happens in all scenarios with a single disk - eventually an operating system will hit a bad sector which it will have to deal with.
In other words, why does the RAID controller crap itself when it can't read (with no recourse at all, according to these articles) when it could just do what every other hard drive does and return 'sector unreadable'. Then the operating system can just remap it etc.
I know in some situations one would want to be notified of any miniscule error, but it should be possible to ignore the warnings.
Sounds like ZFS's proactive sector sweeps across all managed drives would handily solve the problem the article raises with conventional RAID.
If you want "proof" i can dump out zpool status/info/log/etc to show i'm not lying. Note the pool is ~50% in use so its not a great example. Also its raidz2 (raid6) so not a direct comparison. I also bought each drive from different lots to hopefully ensure if a drive failed i'd have 2ish days to get a replacement.
Checksumming filesystems let you find the faulty drive by reconstructing data from each n-choose-(n-1) drive set and finding the set with the correct hash.
Filesystems using FEC instead of raid (5/z1, 6/z2 ...) can also correct data errors, but I'm not aware of any consumer-level filesystems that implement it. I'm not sure why. Doesn't Amazon use it for S3? Data block and FEC data layout on a disk array has to be a solved problem.