RAID's Days May Be Numbered
enterprisestorageforum.com
enterprisestorageforum.com
I am concerned that no one in the open source world appears to be working on chunk RAID. We may end up in a situation where it's fundamentally unsafe to run Linux/BSD software RAID on our 4 TB disks.
Considering the trends in storage capacity, I think the mirroring approach done with ZFS is a clever solution. One could even conceive a file system that would position multiple copies of blocks of frequently accessed files all over the disks in order to maximize throughput, keeping only a minimum redundancy condition declared by file.
I am happy file systems are no longer boring ;-)
The ECC idea could work, but it would be best to build it into the drive's firmware. This way the drive would get read errors less often, could attempt to rewrite the block or to relocate it somewhere else in case of a "bad enough" error.
I think the fix will come from the filesystem. A filesystem that automatically mirrors (or distributes ecc code for) files, and checks checksums is the future.
As soon as a disk fails and is taken out of the pool, the file system could (conceivably, at least) start replicating all data on the remaining disks maintaining the redundancy minimum and shrinking the window of vulnerability to multi-drive failures.
By the time a second disk fails, all data on it could be replicated elsewhere and all you would observe would be a shrinking file system (and urgent messages from the server).
If we want the same capability at a layer below the file system, say the raid controller, than we end up needing a much more complicated implementation that is in effect a layering of two filesystems. Given general trends towards clustered commodity hardware I'd expect the software to keep getting smarter and the block level devices to keep getting dumber.
It seems like there's a role that's missing in the storage hierarchy. Something like an extent manager that doesn't maintain all the actual metadata, leaving that to the file system above, but that does have a concept of a tree of indirectly referencing extents.
Something similar to how allocation and garbage collection can be provided by libraries in c++.
A thought that occurred to me is I don't know anyone that replaces a hard drive when it gets too old. They replace them when SMART says failure is imminent or the drive has already crashed. Hard drives should be treated like tapes. No one replaces tape when they are too ragged to be used. Tapes are replaced on a scheduled (like 30-40 writes). Replace drives after a duty cycle of 5,000 hours and you can minimize your exposure.
Personally, I would quadruple that:
http://usenix.org/events/fast07/tech/schroeder/schroeder_htm... http://usenix.org/events/fast07/tech/full_papers/pinheiro/pi...
They same way tire manufacturers rate their tires for so many miles or kilometers. No logical person waits for a car tire to self destruct before replacement. You note the odometer and make a record of when to change them. So you're not the unfortunate schmuck stuck in snow storm with shredded radials.
Data storage is the antithesis of good maintenance. Instead of proaction its reaction. You buy and install a drive, wait for it to die, then replace it.
The likelihood of a drive failing does not linearly increase with it's age. For relatively young drives, the opposite is true -- if the drive has survived a few weeks of heavy use, it is much LESS likely to fail in the future than a fresh drive off a shelf. Tires are changed routinely because they wear -- hard drives don't so much wear as spontaneously fail. Changing them "before failure" costs you money and increases the amount of drives that blow up in your face.
Unless something fundamental has changed about hard drives I think my point is still correct. Hard drives are electromechanical devices that also wear; bearings wear, fluids evaporate, and heads hit the platter. They spontaneously fail like light bulbs spontaneously fail. It's not spontaneous at all. If you know a light bulb is about to blow or is near the end of it's life cycle you change it.
In the greater scheme of things the cost of the hardware is miniscule compared to the data on it. If drives start randomly blowing up in your face it's time to get a different model.
What I'm getting at is we put more care into the maintenance of our cars than we do our data. If you think hard drives are expensive than what do you think of the cost of productivity while an office full of people wait for a RAID to rebuild. Instead of waiting for the inevitable failure wouldn't it be better to cycle old drives on a friday evening when usage is low.
RAID6 is just a bandaid on a bandaid. It came about because drives are of the same age when a RAID is built. And start failing around the same time. RAID 5 can recover from one failed drive, RAID6 can do 2. But you are still vulnerable if a 3rd. But this is still reactive thinking, "I will replace them as they fail", rather than using proactive solutions.
References, please.
You say that like it has something to do with the article. It doesn't and it is not implied in any way.
RAID is a useful backup for drive failures, if you're able to replace failed drives within a certain amount of time, but not for manual deletions. DVDs and tapes are useful backups for manual deletions, but not for fires. DVDs and tapes, stored at an offsite location, are useful backups for manual deletions and fires, but possibly not for earthquakes.
Since all of the different backup and archival methods have major pros and cons, the only way to say something meaningful about how useful they are is to talk about them within context.
RAID is a redundancy mechanism. It's there in the name. rdiff-backup is a backup tool.
I don't know how to do it on Windows. There's probably a GUI. Although it seems that smartmontools has Windows support.
So, take SMART your data with grains of salt and all that.