Has RAID5 stopped working? (2013)
zdnet.com
zdnet.com
There are also lots of ways to mitigate this problem even with a high error rate. In particular, if you get a bad sector, don't just kick the disk immediately out of the RAID; try recovering that one sector first, using the other disks in the RAID. The drive's built-in sector remapping should then bypass the newly-bad sector.
I wrote about this in detail in 2008: http://apenwarr.ca/log/?m=200809#08
In addition, the reading into the BER and the immediate conclusion that RAID stands no chance is not quite correct. If you did your disk scrubs your disks are mostly clean and the likelyhood of failure is not that great. I was responsible for tens of thousands of disks in a storage system that did do disk scrubs and it did find occasionally a bad sector[1] and I've never even once seen a double failure as is often warned about in the "RAID is dead" articles.
Please do your disk scrubs to avoid sectors falling off and remember that RAID is not backup, extra redundancy is still needed if you really care for your data.
[1] I didn't have the sense of mind to compare the rate of error detection to the official BER.
So here, they're saying that if you have a RAID 5, and one disk fails, then when it has such an error, it's suddenly game over for the entire array - the rest of the data on the drive must be copied to a new array, because the firmware says so. Apparently that's why RAID 5 is a bad idea.
Surely the same argument can be applied to hard drives on their own? If Windows checksummed every file before using it, and forced copying every byte onto a new hard drive whenever an error was found, hard drives themselves would stop working as a meaningful method of storage in 2009. But it doesn't, you (likely) get some garbled text, and it's usually not a problem. File systems get corrupted etc etc (unless you use something like ZFS), and RAID 5 handling it in a particularly non-graceful way is just an implementation detail.
So really, this is just a problem with the controller firmwares forcing you to make a whole new array. It's not a problem with RAID 5.
Or have I missed something?
I know that RAID 5's not a good choice in this modern age (cloud, big SANs etc, SSDs etc), but surely the 'a sector will be bad, and you will not be able to rebuild after a failure' problem would be easily fixed by the raid controller company if their clients complained enough?
> So the read fails. And when that happens, you are one unhappy camper. The message "we can't read this RAID volume" travels up the chain of command until an error message is presented on the screen. 12 TB of your carefully protected - you thought! - data is gone. Oh, you didn't back it up to tape? Bummer!
So, at this point, we've got two hard drives containing millions of sectors for which one or two are bad, and the software claims that the whole thing is broken, and we have to find a new array?
As far as I know, RAID 5 is block level - the block level failure should only destroy one block. All of the others (apart from the other dead ones) are fine. This sort of thing happens in all scenarios with a single disk - eventually an operating system will hit a bad sector which it will have to deal with.
In other words, why does the RAID controller crap itself when it can't read (with no recourse at all, according to these articles) when it could just do what every other hard drive does and return 'sector unreadable'. Then the operating system can just remap it etc.
I know in some situations one would want to be notified of any miniscule error, but it should be possible to ignore the warnings.
Sounds like ZFS's proactive sector sweeps across all managed drives would handily solve the problem the article raises with conventional RAID.
If you want "proof" i can dump out zpool status/info/log/etc to show i'm not lying. Note the pool is ~50% in use so its not a great example. Also its raidz2 (raid6) so not a direct comparison. I also bought each drive from different lots to hopefully ensure if a drive failed i'd have 2ish days to get a replacement.
Checksumming filesystems let you find the faulty drive by reconstructing data from each n-choose-(n-1) drive set and finding the set with the correct hash.
Filesystems using FEC instead of raid (5/z1, 6/z2 ...) can also correct data errors, but I'm not aware of any consumer-level filesystems that implement it. I'm not sure why. Doesn't Amazon use it for S3? Data block and FEC data layout on a disk array has to be a solved problem.
The most often cited issues is with a full disk failure and then having a sector already dead in another disk that you weren't aware of. Do disk scrubs and you will mostly avoid this risk. Your array (software or hardware) should be capable of disk scrubs, most if not all are capable and will do it.
[1] There is a failure mode whereby frequent writes to one sector/track will cause damage to nearby sectors in which case if your RAID stripes are static you may have multiple sector failures in multiple disks in the same stripe. It does require hundreds or thousands of rewrites for all I know so your workload needs to be extreme to reach that case.
- Consumer devices - The later Buffalo terastations (Consumer NAS) and other 'low' price devices for some stupid reason tied the RAID array to their own hardware (like firmware) if the unit fails, the array fails.
- The size and 'time' to rebuild places a great stress on disks, therefore causing a potential 2nd drive loss. Rebuilding HDs for 250GB-500GB took a 90 min or so, a 2TB rebuilt takes several hours, all HDs are grinding to recover data at the same time, for hours as well, and this is when you are likely to lose more than the 1 drive and your array.
- Consumer and SMB NAS devices tend to buy the hard drives from the same manufacturer, at the same time, so when one drive fails, the other drives are just as likely to fail in the same time period. I lost two 1TB backups when the HDs had a 'suicide pact' and when I opened the Hard drives, I saw they were literally made the same day. To be extra paranoid, I used to have 2 or 3 different brand of HDs when I had my NAS. - I recommend Synology units and constantly monitor the health of your hard drives. Totally agreed that RAID1 should be the golden standard. I've used several brands and Synology was rock solid.
So while folks who really do need to reliably store 500GB of stuff might still be looking at RAID, for the 'average' end-user they have a far better solution than complicated multi-disk setups.
Imagine a typical American or Canadian "broadband" consumer internet connection with 10Mbps down and 1Mbps up. If you keep your computer on 8 hours a day, it will take just over 9 months to upload 1TB of data to the cloud at that speed. A multi-TB RAID array would take years to back up. (This is why I think Backblaze et al. can afford to offer "unlimited backup" for a flat monthly fee. The customer's ISP is doing all the limiting for them!)
It was a major PITA when I first started using online backup services while in Canada. (TekSavvy FTW!) Despite paying for one of their faster plans, I had to keep my computer on for a couple of weeks 24/7 to make the initial backup.
'YOUR PHOTOS BACKED UP TO CLOUD WHILE U WAIT'
As soon as you get home you can access your backup (transparently) from the mall node. Then when the drive gets back to base it does another verification run and instructs the node that it can re-use that space.
What's also fun: USB speeds. I found a transfer time calculator - http://techinternets.com/copy_calc - and, well, 4 terabytes at USB 2.0 speeds (480 mb/s) would take over 21 hours, or 3 hours at USB 3.0 speeds. Those are at optimal wire speeds, less 10% for overhead.
If you use RAID5 on any size array with out making backups you will lose data. Does he really think that with the invention of 1TB+ drives that it's the first time two drives have failed at once?
In the age of really fast SSDs Raid 1 might be the better choice though, unless you really need a huge amount of fast storage.