RAID5 definitely has a purpose, notable small arrays where RAID6 would be wasting space or if you have nodes to failover too (ie, you have 3 nodes, run each with raid 5, if one goes belly up during a rebuild you still have 2 left and you can reprovision the third).
RAID6 would likely be at the limit of most modern implementations so you'd have to jump to (n,n-3) RAIDs or higher. Of course there is always RAID1(0) if you like 50% storage efficiency.
This doesn't quite make sense to me. Why would the size of the array make any difference as to whether or not the extra drives needed by RAID6 can be characterized as waste?
> or if you have nodes to failover too
In essence, that's RAID5+1, which makes me wonder: If you have node-level redundancy, why bother with the intra-node redundancy?
> RAID6 would likely be at the limit of most modern implementations so you'd have to jump to (n,n-3) RAIDs or higher.
It's not clear to me what limit you're referring to, but, considering how long RAID6 has been implemented, especially in hardware, and how far computing power has increased since then, such a claim is dubious, if not extraordinary.
RAID risk is about your rebuild failure rate, in this case URE (Unrecoverable Read Error). If one such is likely to occur during your rebuild, that can lead to a lot of problems.
In RAID5 until you hit about 2 TB disks and up to about 4 or 5 disks total the risk of a URE during the full read of the array is fairly low. Once you go above that you risk loosing data to URE errors.
In RAID6 you essentially multiply the URE rate together, which means you have a much much lower error floor and you can repair bigger arrays.
In a cost-benefit analysis this means that if your array is small enough, a RAID5 gives you more effective disk space with little additional risk.
>If you have node-level redundancy, why bother with the intra-node redundancy?
A failover is still a failover and can reduce performance and it reduces your remaining failover margin. You want to keep your failover rate low, though if you doN't particularly care you can use RAID 0 too.
>It's not clear to me what limit you're referring to, but, considering how long RAID6 has been implemented, especially in hardware, and how far computing power has increased since then, such a claim is dubious, if not extraordinary.
This is not a CPU power limit, rather RAID6 will in the next few years hit the spot where the risk of loosing 2 drives during a rebuild, due to the large drive sizes, becomes large (URE is still low IIRC), so at that point you want more redundancy to keep the rebuild risk low.
This seems to be the keystone to your reasoning, and I'm not sure it's true, in practice. IME and in the literature I've read, modern RAIDs operate assuming a whole-disk failure (both on write and on read). Sub-whole-device errors seem to be an issue of concern limited to ECC in RAM. Do you have pointers to literature about this in the context of RAID?
> A failover is still a failover and can reduce performance and it reduces your remaining failover margin.
I think that means that the node-level redundancy isn't actually equivalent to RAID1-style redundancy (where there's no explicit failover), which is important to mention.
Of course, my question wasn't really "why bother" so much as asking if that architecture isn't far more wasteful than, say, RAID1 (or even just RAID6/RAIDZ3) intra-node and less node-level redundancy.
> You want to keep your failover rate low
I'm pretty sure best practices are the opposite of this. One can only know if the redundant copy actually works by using it, so a system that can failover more often is better. Ideally, the system is, like RAID1, requires no explicit failover and merely load balances both copies.
> RAID6 will in the next few years hit the spot where the risk of loosing 2 drives during a rebuild, due to the large drive sizes, becomes large
You'll need to quantify these, though, since, though drive capacities have continued to grow faster than transfer speeds (for mechanical drives), it's still in the OOM of 30 hours for minimum rebuild time. What's the risk of a second and third drive failure in 300 hours (1/10th the maximum rebuild rate)?
URE Rate is documented in most handbooks or manuals on harddisks, it's usually about 10^-12, though this is a worst case error rate and -13 or -14 is more realistic. At that rate you means your harddrive will return a read error every terabyte to every hundred terabytes.
A URE can cause the raid controller to think the HDD is defective and crash the array if no parity is left to compensate.
>Of course, my question wasn't really "why bother" so much as asking if that architecture isn't far more wasteful than, say, RAID1 (or even just RAID6/RAIDZ3) intra-node and less node-level redundancy.
Depends on your use case.
>I'm pretty sure best practices are the opposite of this.
Best practise is to test your failover regularly, correct, but this doesn't mean you shouldn't minimize your actual failover rate.
>What's the risk of a second and third drive failure in 300 hours (1/10th the maximum rebuild rate)?
It's not that low, especially since a lot of arrays contain similar enough harddrives that concurrent failure is possible (although you can make it less likely by avoiding similar batches). And again, that URE rate has a risk of marking of your array as dead if you get unlucky.
The kind of literature I'm looking for is based on real world data, since spec sheets are generally marketing documents and therefore, at best, uninformative. This is particularly noticeable when a single number is used for a statistic, as in this situation.
Let's take a current datasheet [1] which actually says 1 sector (unclear if that's 512 or 4k, but let's assume 512) per 10E15, max, a full 3 OOMs higher than what you mentioned. That's 512PB, which, at the highest listed transfer rate, would take almost 545k hours, corresponding to slightly above a 1.6% AFR for worst-case URE-caused failures.
Since I don't believe the spec sheet's 0.35% (average) AFR, but since real-world AFRs are in the low single-percent, I'm staying with the conclusion that read errors need not be considered separately from any other underlying cause of drive failure.
> Depends on your use case.
I'm not convinced that "waste" can ever depend on use case, when examined under the narrow lense of redundancy. Regardless, it was a use case you, specifically brought up, so the question remains open.
> Best practise is to test your failover regularly, correct
You're misunderstanding my assertion, which wasn't about regularity but frequency. I made no statement as to regularity (and it may even be better, with discrete failovers, to do so irregularly, rather than regularly).
My assertion is that best practice dictates frequent failover.
> It's not that low
> risk of marking of your array as dead if you get unlucky.
You still haven't actually quantified the risk. "Unlucky" and "possible" are too hand-wavy to engineer around.
Quantifying the risk, even if it's a very coarse approximation, is a prerequisite to avoiding FUD-based decisions/actions. Usually, the most visible consequence of the latter is the waste alluded to earlier, manifesting as higher cost and/or lower performance. The less visible consequence is misallocated resources/attention from something relatively likely (e.g. human error causing catastrophic data loss) to something much less likely (e.g. single sector read error causing catastrophic data loss).
[1] https://www.seagate.com/www-content/datasheets/pdfs/exos-x-1...
Of a high profile and high quality enterprise disk, unlikely to appear in 99% of RAID setups. Congratulations.
Most RAID setups in the wild consist of low quality enterprise or even desktop platters, high quality setups are rare and expensive.
Esp. considering that lots of consumers have a NAS or MyCloud/eqv. at home and SMB doesn't buy those harddrives either, I don't think this is a valid comparison.
[https://www.seagate.com/files/www-content/product-content/na...]
[https://www.seagate.com/staticfiles/docs/pdf/datasheet/disc/...]
These drives here have an OOM higher error rate on the data sheet and they're still not very common, plus the error rate is given in bits here, shaving of 3 or 5 OOM again, calculating it out gives you an bad bit every 12 TB or so. Depending on how your RAID works, this can ruin the entire stripe or the sector it gets the hit on.
12TB isn't a lot and there is a good probability that a 3x4TB RAID5 array will hit such an error and a 3x2TB RAID5 has a 50/50 chance of having it during a rebuild.
>I'm not convinced that "waste" can ever depend on use case
It definitely can. Consider the use case of "office documents" vs "high performance video storage". With office documents a high throughput is not needed so we can reduce storage costs by using a RAID5 or 6 depending on array size. Video storage for example for CCTV requires more throughput an a RAID10 will give you better performance at some lost effective storage.
You can of course put your office documents on a RAID10 but that is the wasted space. It is not necessary to maximize the RAID performance and the low storage needed for office documents means you can use RAID5 on small arrays fairly safely, RAID6 if you need more.
Waste is a matter of efficiency in all domains, in this case means picking the RAID level that maximizes storage efficiency, performance goals met and safety as best as possible. Picking one that trades off one for the other is waste or even dangerous.
>My assertion is that best practice dictates frequent failover.
I don't think I ever met a sysadmin that insistent frequent failover due to failure is considered best practise. Regular testing yes, regular failure, no. Failure is expensive, testing not.
>You still haven't actually quantified the risk. "Unlucky" and "possible" are too hand-wavy to engineer around.
They're not really, plus you can easily quantify risks as you have done using datasheets. Though the moment your array isn't a singular batch of exactly the same harddrive these estimations become a lot harder.
You don't need to exactly quantify the risk, it is sufficient to know the margins of the risk or even if you are going into the margins of the risk, ie a risk model.
I don't need to know the exact and perfect probability of a read error on the array, I need to know if the array is likely to experience one (with likely being >1% or any other >X% during a rebuild or during normal operation). This question is more easily answered and doesn't require more than estimations.
This trades cost efficiency for safety margins meaning that an array has a comfortably low actualized risk of failing. That is what I consider actual engineering there.
> Most RAID setups in the wild consist of low quality enterprise or even desktop platters, high quality setups are rare and expensive
> they're still not very common
More extraordinary claims, requiring extraordinary evidence. Even Baraccuda Pro models (which do cost more but not prohibitively so) have the same read error rate, but I didn't choose that datasheet because it had less data, overall.
> Esp. considering that lots of consumers have
I call "red herring" on this, since consumers also aren't going to have the kinds of choices we're discussing, nor read these forums, nor the original article (which is the context for this whole discussion).
> the error rate is given in bits here, shaving of 3 or 5 OOM again
I'm not convinced of that, since the tables look identical between the drives. Maybe it's a sneaky marketing ploy, but maybe not. Ultimately, you need real world data, which you consistently haven't provided.
Absent that, it seems as though you're relying on assumptions, and my original conclusion, based on the data that has been published by the likes of Google and Backblaze, stands.
> Waste is a matter of efficiency in all domains
You seem to have re-defined waste, so I can't really speak to it.
> I don't think I ever met a sysadmin that insistent frequent failover due to failure is considered best practise.
I fear you are, again, misunderstanding. I didn't mention failure as a cause, although the term "failover" could lend itself to confusion, with the substring "fail" being in it. I used the term only in the sense of "switchover".
Perhaps I merely misunderstood you originally. You did initially state "A failover is still a failover and can reduce performance" which, even assuming you meant failover-due-to-failure, the assertion is questionable in the context of justifying node-level redundancy on top of RAID-level redundancy, if switchover (not due to failure) is engineered to be frequent (or even continuous).
> you can easily quantify risks as you have done using datasheets
I'm not agreeing or disagreeing as to its ease, but I'm asking you to go ahead and perform this, which you assert the ease of, since that seems to be the basis of your point.
(Earlier, I just made some single-disk calculations based on the spec sheet, not array risks.)
> I need to know if the array is likely to experience one (with likely being >1% or any other >X% during a rebuild or during normal operation). This question is more easily answered and doesn't require more than estimations.
Agreed. As I mentioned, coarse estimates (if based on real data) are plenty good enough. However, even a coarse estimate assigns some number to it.
Given your question above, what's the answer, for likely-being->1%, during a 300-hour rebuild? How did you arrive at that answer?
If you're getting performance benefits from the RAID1, then the extra drives may not be wasted, but that's a separate topic.
[1] Which I argue is a misnomer. To me, "warm spare" makes more sense, since it's powered up but not actively synced in any way. Some systems, configurably, even spin down such spares.
[2] As I mentioned in https://news.ycombinator.com/item?id=17855632 some failures are power-on/spun-up related, which makes the availability of a truly cold spare even more beneficial.
I thought zfs was not incompatible with RAID. How else would you run it out of curiosity?
- An external Drive. (which is only connected to copy files)
- An external cloud service.
- Another computer that you have.
You should probably do a couple of these.
I personally backup home computers (using borg) to a home server, that server has a 2.5 2TB external HDD connected to it (2 other 2.5 external drives are kept outside of the house). A backup of important files from the nas (including the computer backups) gets copied over to the external drive nightly. Weekly the drives gets rotated.
The really important stuff is also backed up offsite on a daily basis.
The way to think about backup and DR is "what would it take to destroy all of this?" and keep making the answer more and more extreme until it is so horrible that if it actually happens you won't care about what you lost.
PS: Also always remember that a backup you don't test isn't actually a backup.
The other "rule" that many people follow is 2 is 1 and 1 is none. The idea is that anything that isn't backed up isn't protected and can't be relied upon to exist.
A cloud service like backblaze, google drive, amazon cloud drive, etc. is a good secondary backup for a lot of people even if it'll take you a month or two to get your data there to begin with.
Disks are large and cheap enough to afford losing 50% of your capacity, and it's much faster while in use and when rebuilding.
Other examples are ZFS's RAID-Z, which can support even more parity for even more resiliency to drive failure.
Contrary to a sibling comment, RAID1+0 does not supersede RAID5, as it has existed at least as long and has always had different trade-offs.
How many drive failures one may need to be able to survive, given historical drive failure rates [1], current drive capacity vs. transfer rates (i.e. minimum rebuild time), and individual parameters (e.g. number of drives per array, acceptable magnitude of performance degradation during rebuild), is left as an exercise to the reader.
As the OC suggested, this number can still be 1 (i.e. RAID5) for SSDs.
[1] optionally including or excluding "black swan" events such as the flooding in Thailand that wiped out thos disk factories, rendering an entire "generation" of HDDs, manufactured in haste elsewhere, far less reliable
Care to elaborate on this? Disk specs? Resilivering time?