Understanding RAID: How performance scales from one disk to eight
arstechnica.com
arstechnica.com
> _We find that with hardware RAID, it's frequently difficult to tell whether you're nuking your array or importing it safely. So in the event of a controller failure and replacement, you may end up sweating bullets, YOLOing, and hoping._
This was the main reason I decided to switch to Linux software RAID from a 3ware card years ago.
A thunderstorm caused a power trip and after powering the system back on, the 3ware controller lost its EEPROM contents that held the PCI vid/pid. Fortunately, it was just a RAID 1 mirror. We used dd to cut past the metadata header and managed to mount the ext2 filesystem.
Using software RAID, I’m confident that we can take the drives to another machine in a pinch and still be able to mount it without much hassle.
In 05 Dell sold me a workstation and internal RAID array that suffered a drive fault requiring replacement and rebuild. Unfortunately the engineer who was handholding me through the steps just to make the changes under our support conditions of the job being theirs and otherwise we left alone, told me it was a LSI card and not a Adaptek installed. This was a very entry level card and the only such card I've experience of. (Accustomed to the luxury of the high end) So I had to take him on his assertion it really was a LSI OEM PERC controller. Erm, not exactly. I learned this when a second connector dislocated and I blithely yanked and reseated U32PSCSI ribbon neatly and promptly as I imagine anyone would. Only the PERC was Adaptek and not LSI and for reasons my recall may be protecting me from* failed the second drive in the R10 array.
I vowed to never run any hardware RAID again and I have maintained my resolution until last month, if specifying a rack of dedicated storage counts for proprietary hardware in equivalece of black box opacity.
I should have written my experiences and blogged them because I took this loss of the subsystem so badly I became obsessive about learning first how to reconstruct the borked array from raw disk, a pleasure indescribable in polite company, and carried on until I was holding forth in charge of a budget for storage space in my dreams I never imagined presiding over. Every which way I believe only in software control of stores. I'm a true disciple to the cause. This r as rack of proprietary storage hardware only got its PO after the methods of code escrow had been thrashed to death in the final thrills of fall and last winter. (Incidentally the outside legal lead wrote me privately afterwards and pondered out loud whether the escrow didn't create a backdoor kind of vendor lockup because of the potential for contaminating employees with IP during emergency contingencies operating. Thanks Buddy! I'm not sounding like it but I'm actually grateful for that consideration if only because it sounds good for avoiding any repetition of the process just endured. It's a dual licensed codebase but being on the commercial code and the open code trailing in develop... I'd love to see what eg DDN routines are like I'm absolutely sure eager audience awaits publications like that.
Surely the lock down must have awoken the industry to the needs of self sufficient repair and diagnosis.
Hardware RAID is just data corruption waiting to happen, simply because there are more vulnerable components and interconnects along the data path.
It's much better to checksum data while it's still protected by ECC RAM and CPU ECC caches – which is something only software RAID can do. Although even that is not a total guarantee.
There's no longer reason to assume a slow SoC on a RAID adapter (even with XOR acceleration) would perform better than an Intel or AMD server CPU, as long as I/O path is not bandwidth limited.
Also hardware RAID allows for a more consistent performance under heavy system load in my experience.
This is not what I was talking about. Of course they do. The problem is that data needs to somehow be transferred to the adapter first, and this alone adds a point of failure. Controller can (and too often do) have hardware and firmware bugs. The less components the lower chances for corruption.
> battery backed-up cache for the past decade, but supercapacitors and flash memory
No difference to software RAID there. For example ZeusRAM and Optane are used for this. Of course high endurance SLC flash with supercaps is also an option.
> Also hardware RAID allows for a more consistent performance under heavy system load in my experience.
I'd tend to count this subjective, but I'll grant you this one, it's at least plausible. I haven't encountered this issue, though – today's hardware has insane bandwidth. SW RAID can saturate 100 Gbps ethernet, how much more performance you really need?
Even if performance was slightly degraded, data integrity is more important.
"640K ought to be enough for anyone". To answer more fully, you need as much as you need. You may not, but others will need more.
Notoriously embedded CPUs have a much worse price/performance ratio than mainstream CPUs so almost inevitably they use a much weaker CPU than is needed.
This is one reason why WD is involved in RISC-V, so they can escape this, and another reason why Intel's storage strategy (ex. optane) has worked out, because the main CPU has power to RAID NVMe SSDs and off-brand CPUs don't.
But yes, not a good post.
Nobody said everyone everywhere.
Nobody said forever.
Something wrong with dynamic disks for software raid for the os in windows?
https://docs.microsoft.com/en-us/windows/win32/fileio/basic-...
Or storage spaces? (the latter i have not used)
https://support.microsoft.com/en-gb/help/12438/windows-10-st...
Remember to enable ReFS integrity streams! ZFS protects full integrity by default. NTFS supports only protecting metadata, file data integrity is not protected.
Fill the volumes with known data. Randomly corrupt some blocks out of disks (but not from both disks at same offsets!).
Observe which solution recovers and returns correct file data and which does not.
Hint: you might want to stop using dynamic disk RAID after this experiment...
Storage Spaces operates on a higher level, building on lower level volumes. I'd feel way more comfortable using it on ReFS, at least once it can be considered sufficiently mature in the future. See: https://docs.microsoft.com/en-us/windows-server/storage/refs....
No qualms about trusting ZFS, it's very solid on good hardware. Of course running ZFS on Windows in production might be a Very Bad Idea, so do it on FreBSD, Solaris derivatives or Linux instead!
Last I checked... VMWare VXSI does not support a software RAID at all anywhere. I think that’s still fairly common in server environments.
See for example https://www.vmware.com/content/dam/digitalmarketing/vmware/e...
You get what you pay for.
I've always wondered whether anyone has gotten fed up with this, and decided to just virtualize Windows Server under Linux in order to feed the Windows Server VM a virtual disk that sits atop all the Linux storage-layer tech (but where otherwise the Windows VM gets all the rest of the computer's resources passed through.)
(Come to think of it, I've considered doing the same for a Hackintosh at some point. It got too contrived, but I'd like to come back to it one day.)
But you're 100% correct. I have friend who have few RAID controllers in his office just in case that one of it's servers fail. And all of them are purchased in same order to be from same batch.
Using this RAID calculator gives me worrying numbers for high capacity arrays, even with enterprise hardware:
https://wintelguy.com/raidmttdl.pl
Also, expanding the capacity of an array is a needlessly complicated process:
https://raid.wiki.kernel.org/index.php/Growing
The best solution seems to be something like Ceph... But that isn't practical for a home lab.
Yes? But that's unrelated to the comment you were replying to.
Please note that the Author Jim Salter also has a nice podcast: https://techsnap.systems
Also interesting: https://arstechnica.com/gadgets/2020/02/how-fast-are-your-di...
My main gripe with the article is that the topic of latency seems underdeveloped.
The testing I've done using some of the parameters discussed on the article shows extreme, probably unrealistic latencies on the storage.
This causes RAID rebuilds to take longer. But I am not aware of any hard evidence that this means that this increases the risk of a failed rebuild.
There are more sectors that could be bad but you need to do a patrol read / scrub at least one month to detect them in advance.
I think there are a lot of scares about RAID of which I really wonder how much of it is rooted in anekdote and folklore.
I had five Seagate 7200.11s (terrible, shitty drives in many ways) in a RAIDZ1 pool. One failed, I started a resilver. Another started failing during that time. The rebuild slowed significantly due to the second failing drive. ZFS resilvered eventually then the second drive failed completely. That was at only 26k hours or so. I think a third started acting up shortly after but I was already replacing them all.
These were only 500GB drives, and the chances aligned on two of them....
I only use RAIDZ2 now (and I don't use Seagates).
Question about ZFS and its RAIDZ: do you have any recommendation (personal experience, links, books, ...) concerning parameters/setup to be used when setting up a RAIDZ(1 and 2)?
I'm new to ZFS and I already had to do a lot of tests when using ZFS only on 1 HDD until I finally managed to get good performance out of it (used by a "Clickhouse" database which itself writes data in a CoW-style => I had to raise the "recordsize" to 2MiB) and I imagine that with RAIDZ it can get more complicated?
I would like to set up a RAIDZ for the database (again, "Clickhouse", which generates multi-GB files) and two more to be used as simple NAS (a main one and a backup, storing files of all sizes).
I searched a lot and found some websites which were "ok", and bought as well 2 tiny books ( https://www.amazon.com/Introducing-ZFS-Linux-Understand-Stor... and https://www.amazon.com/ZFS-Linux-Administration-William-Spei... ) but the books were mediocre and the informations I found in the web were sparse and a bit oldish.
Cheers
I'll try to think of some others.
Honestly the best way is to start tinkering and see how it works for you as it can depend heavily on what you're using it for.
But yes there is a lot of FUD out there.
One thing to beware of, RAIDZ1/2 is limited to the IOPS of one drive per VDEV... Can be quite limiting for some things.
Didn't know that website - it has a lot of interesting stuff.
Ok, I'll then do some tests and will see how the RAIDZ behaves.
libeatmydata is a small LD_PRELOAD library designed to (transparently) disable fsync (and friends, like open(O_SYNC)). This has two side-effects: making software that writes data safely to disk a lot quicker and making this software no longer crash safe.
DO NOT use libeatmydata on software where you care about what it stores. It's called libEAT-MY-DATA for a reason.
Also just some latin pedantry, the article mentions the phrase: Caveat imperator a couple of times. I assume to mean Let the buyer beware, which as far as I know is Caveat emptor.
Google latin translating: https://translate.google.com/#view=home&op=translate&sl=la&t...
Does caveat imperator have another meaning (perhaps the emperor is infallible)?
https://arstechnica.com/information-technology/2020/04/under...
He did mention it as a side note, but it deserves more attention I think.
Most consumer drives are still sold rated as <10^14 bits read per error. That's 12.5 terabytes, so in the worst case you could end up in situations where — on average — you are unable read a full 16TB drive without an error. Needless to say this is less than optimal for rebuilding a failed RAID array.
Anecdotal evidence (i.e. very low error rates during ZFS scrubs) suggests that manufacturers underrate their drives and they are much more reliable than that, but it is something to keep in mind.
Fortunately drive capacity completely outpaced my needs for personal data storage in the recent years, so I am happy with JBOD or RAID1, with backups of course.
[1]http://www.raidtips.com/raid5-ure.aspx [2]https://magj.github.io/raid-failure/
I've seen plenty of drives returning invalid data with correct CRC over the years. On reliable server-grade Xeon + ECC hardware. Of course the vast majority of drives never do it, there's just no way to know which ones do until it happens.
Firmware bugs in weird corner cases? Cosmic rays? Perhaps, but I think it's more reasonable just consider it one of those weird things that occasionally Just Happen (TM) and just need to be protected against at a higher level.
All drives produced in the last 30 years or so are running ever more complicated software stacks. For example they all have features that move data at risk to safer locations without the host system requesting or even knowing about it. Their physical (like bits on spinning rust or NAND flash block) and logical (what the host sees) data representations can be completely different.
Plain old CRC errors, though, are way more frequent. I feel much more comfortable about those errors, at least the drive knows the data is corrupted.
I switched to ZFS a few years ago and it's really quite amazing. In fairness, raidz does suffer from the same "resilvering takes a long time and increases the chance of subsequent failures" problem, which is one of the reasons why mirrors are more commonly used. Though unlike RAID, ZFS knows which data on the disk is actually used so can reduce resilver times pretty drastically.
Which, to be fair, makes sense with the subject of the article limiting us to a discussion of 8 drives.
Plus, you’re talking about RAID as a backup, it is not. RAID is redundancy, not disaster recovery. If your backup solution is bas d on your RAID controller series, you’ve already lost.
So, here goes...
The management tooling for hardware RAID is a pain to deal with. Dell's OpenManage tools aren't that bad (other than being stupidly bloated), but every time I have to use LSI's MegaCLI utility, I end up wanting to shoot myself. (Seriously, do a Google search for "MegaCLI cheat sheet". It's _bad_.) For the cheaper LSI FusionMPT-based cards, I've completely given up trying to manage them from the OS, because I've never been able to get it to work reliably. And because hardware RAID controller manufacturers have changed ownership so many times over the past decade, even _obtaining_ the software can be a challenge. And that's assuming that it supports your OS, which isn't guaranteed.
Monitoring the health of hardware RAID arrays is challenging. Not only do you generally need to use a proprietary tool of varying quality to do so (with all the issues stated above), the information you get is typically not that helpful. Most RAID controllers will at least tell you if the array is degraded or not (i.e., one or more disks are missing), but the quantity and quality of that information can vary. Additionally, monitoring the health of the underlying physical disks that make up the array is also important, but getting this information from a hardware RAID controller is tedious (and sometimes impossible), which can leave you blind to the actual health of your storage. The disk health typically reported by RAID controllers (the greem/amber blinky light that indicates a failed or failing disk) is based entirely on the disks reported SMART health check status, which is so prone to false negatives that it may as well not even exist. Oh, and if you're using an OEM RAID controller with third-party disks, you often have to deal with a nag message informing you that the disks aren't "official," which at best is annoying, and at worst can hide legitimate disk issues.
Hardware RAID can have annoying limitations. Want to mix a SAS and SATA disk in the same array? Most controllers won't support that. Want to mix RAID levels on the same disks, like in the case of a distributed database where you might want some redundancy for the OS volume, but the actual database volume can be RAID 0 for performance/space efficiency? Most controllers don't support that, either. Want to create multiple distinct logical volumes from a single array (e.g., a 20 GB volume for OS, and a 10+ TB volume for data)? That's usually possible, but in the case of many LSI MegaRAID controllers (including their Dell PERC rebrands), that can prevent you from expanding the space on the array without a rebuild. Want things like RAID 6, more advanced nested RAID levels like RAID 50 or 60, or tiered storage? If the controller supports it at all, it's probably a premium-feature that costs extra to unlock.
Hardware RAID is slow. The onboard processors on most controllers have no problem keeping up with mechanical disks, but struggle with the performance potential of SSDs. The centralized cache on the controller can be useful in some situations, but the typical configuration is to use this cache while disabling the cache located on the disks themselves, which runs into scalability issues for obvious reasons. And lastly, you're ultimately limited by the throughput of the bus to which the controller's attached, which means that a hardware PCIe RAID controller can't possibly compete with NVMe disks that are directly attached to the processor via the PCIe bus.
Software RAID (at least on Linux) is fast, simple, easy to use, the UX is consistent no matter the underlying hardware, it's well integrated with many existing monitoring tools, and frankly, it's been more reliable than hardware RAID. All of our newer machines use software RAID with either NVMe disks or SAS HBAs, and as our older machines are being repurposed, they are being converted to software RAID where possible. It's made my life a lot easier, and nothing of value has been lost.
Hardware RAID is dead. The only reason I'd ever use hardware RAID again is if I had to use an OS which didn't have decent software RAID support.
But, again, this is very low end to middle end 'hardware raid'. It's single-digit disk stuff. It's a cheapo PCI card in a cheapo server.
You can multiplex hundreds of disks across multiple SAS backplanes. And, yeah, those software RAID [0] systems are really cheap. :-)
Software RAID is deployed in low end systems, sure. But also in very high end systems where compromises can't be made. And anywhere in between.
[0]: For example, see: https://www.oracle.com/storage/nas/zs7-2/
Even hardware RAID is somewhat of a misnomer. Just like the disks themselves, it's just an embedded CPU, varying degrees of hardware acceleration, possibly battery/ultracap backup and a big software stack.
Software RAID can achieve high capacity and/or performance. No difference in principle to hardware RAID there. You can use multiplexing for huge SAS backplanes in either case.
Question for you smart people... is there a reason to use such an old Linux kernel version? They're testing on 4.15, which was released over 20 months ago. Would there be any appreciable difference on 5.6.4 or something in the 5.x range? Perhaps not since hard drives can't really improve much, but I'm curious.
The block layer has been able to deliver very close to 100% of the drives' stated bandwidth for many years now.
Can anyone else weigh in on this claim? In the two jobs where I've been responsible for such equipment, monitoring both the performance and battery status was pretty critical and done with due care and focus.
What it offers (in parity mode) is improved reliability and more chance that your server will be online if your drive dies. In majority of cases, it saves you from restoring from backup. Just replace the bad drive, and away you go! However, never assume that it will protect you from data loss. Backup. Backup. Backup!
That said, an argument such as "system is down while we restore from backup; some recent data may be lost" will not go over as well as "the system remains available, but performance will be negatively impacted while the array rebuilds; no data has been lost at this time".
In the former, you're in crisis mode. In the latter, you'll probably be under some stress, but you're still in the clear.
Edit: But we're all on the same page when it comes to backups. They're a necessity.
I.e. ceph with pods distributed over a bunch of disks without a raid
Not if you need speed it doesn’t.
A tape drive costs ten times what a hard drive with the same capacity costs, and then you have to buy the tapes.
I've been able to read data written to a hard drive 99.99% of the time but I can't say that for tapes. Often it has been human error, but the fact that it's a strange thing you don't do everyday and isn't part of your workflow is a reliability problem that is hard to overcome. (e.g. why a Toyota Corolla is more reliable than a Ferrarri Testarosa.)
In theory you could do offline backup to "the cloud" but probably your net isn't fast enough and the other people at home will be wondering why Zoom isn't working for them.
Recovery is a bit funny with you marking all remaining drives as faulty then plugging backup for resilvering. (Or you can use it directly, less safely, as a normal separate array.)
Must be disconnected and unpowered, yet regularly checked. Not super useful for extreme long term storage but neither are tapes.
Mmhh, sounds weird - what do you mean with "disconnected RAID1"?
Sounds like a misuse of RAID features. What could possibly go wong? Eg. Reconnect the drive accidentally forgetting to mark others as bad and bye bye "backup"?
Also, RAID controllers die sometimes. It's not pretty when that happens, which means you really need a backup of the whole array anyways. So technically RAID is not a backup because you are still not protected from that type of failure.
This seems a bit like someone saying "my blog never needed Kubernetes". Of course not, that's not what it's for.
Obviously those things can be worked around with complete re-architecting of, well, everything, but I'm considering that out-of-scope.
10 is 1+0
I'm using two SATA SSDs in md RAID0 for VMs on my desktop (where each VM gets its own LV). Dangerous thing, but if one of them fails, nothing happens, I will shrug and reinstall.
You can take as an example the Recovery Point Objective and Recovery Time Objective terms used on disaster recovery. Depending on: a) how often you do backups, b) how many data are you backing up, and c) how fast are your network and transfer medium, you can calculate how much data you may lose in case of failure and how much time you'll use on recovery.
- If a drive fails, the system is offline until you restore from backup. Fine for personal use, but not okay if customers are paying you for a service.
- If a drive fails, any changes since the last backup are lost. Fine for long-term archival, not so great for a bank ledger.
- No capability to detect data corruption. RAID with parity can run a scrub to find and repair bit rot. If data is corrupted on your RAID 0, you probably won't notice, and you'll back up the corrupted. If your filesystem has checksums you will be able to detect the corruption but not repair it. Your filesystem probably doesn't have checksums.
If you frequently run I/O intensive workloads, the extra performance might be worth the tradeoffs. If you're looking for your PC to feel slightly faster it seems foolish.
- I've got 1 HDD failure on RAID10 4 x 10TB HDDs, I had to take the system down to do RAID10 rebuild. because running server with huge I/O slowdown the rebuild and I had my fears of second HDD failure. (I had to take the system down anyway). RAID didn't help in my usage scenario.
- I guess deduplication fixes this already.
- Any advice on ways fixing this and preserving my RAID0 usage scenario?
I manage a small (<100TB) storage system for our office. We have servers using RAID10 SSDs and we have other servers using RAID10 10k SAS HDDs. The read performance on the SATA SSDs is great but the writes are atrocious.
If a 10K SAS HDD array is outperforming even a single SATA SSD, either your I/O access pattern is extremely sequential, or there's something wrong with how your storage is configured.