I think the fix will come from the filesystem. A filesystem that automatically mirrors (or distributes ecc code for) files, and checks checksums is the future.
I think the fix will come from the filesystem. A filesystem that automatically mirrors (or distributes ecc code for) files, and checks checksums is the future.
As soon as a disk fails and is taken out of the pool, the file system could (conceivably, at least) start replicating all data on the remaining disks maintaining the redundancy minimum and shrinking the window of vulnerability to multi-drive failures.
By the time a second disk fails, all data on it could be replicated elsewhere and all you would observe would be a shrinking file system (and urgent messages from the server).
If we want the same capability at a layer below the file system, say the raid controller, than we end up needing a much more complicated implementation that is in effect a layering of two filesystems. Given general trends towards clustered commodity hardware I'd expect the software to keep getting smarter and the block level devices to keep getting dumber.
It seems like there's a role that's missing in the storage hierarchy. Something like an extent manager that doesn't maintain all the actual metadata, leaving that to the file system above, but that does have a concept of a tree of indirectly referencing extents.
Something similar to how allocation and garbage collection can be provided by libraries in c++.