There's tons of OSS to manage all layers of this stuff now.
There's tons of OSS to manage all layers of this stuff now.
RAID5 gives you really the worst of all worlds. You get poor performance (n/4) due to the number of operations required and you get the fragility of a RAID storm that comes from a cascade RAID failure when a drive dies. A "stripe of mirrors" (RAID10) gives you the safest and the highest performance array (at scales where these decisions are being made which most people would see anyway). The common argument against is the cost (usableCapacity = totalCapacity/2) but my counter argument is that RAID5 is a timebomb with a secret countdown. RAID5 is great if you don't need the data on the array.
Source: >20 years running complex large arrays
But I do agree with you, it does make the problem significantly easier/different.
ZFS and Ceph are marvelous creations.
What are your thoughts on implementing RAID 0 over four RAID 1 arrays, each consisting of three disks, resulting in a total of 12 drives?
S3 keeps multiple copies in multiple geographic locations.
Having more locations is actually helpful for this. Because of the independence of different locations, the risk of data loss is in some ways lower than a local RAID system where all drives might be fried at the same time by a single event.
50-ish drives in a single 4U server is done.
or, if you go fibre channel, just replicate the writes across multiple storage systems and be done with it. (this requires a fibre channel network though, which is neither easy nor cheap).