In other news, all of the WDRed CMR drives in my ZFS array started throwing read errors simultaneously so I've ripped them all out.
I'm about done with WD.
In other news, all of the WDRed CMR drives in my ZFS array started throwing read errors simultaneously so I've ripped them all out.
I'm about done with WD.
I don't get it, is the implication that them showing errors at time should be a freak occurrence, and such an occurrence means that WD drives are low quality? Drive failures are more correlated than you think. If you bought and installed the drives at the same time, such a failure case is totally expected. The drives probably came from the same batch, and experienced similar amount of wear, so it's only logical that they experienced the same failure case.
They'll still fail, but hopefully they'd be less likely to do it at the same time.
But that means you need RAIDZ2 or erasure coding to be reasonably safe, which takes you well outside what most of these turnkey systems can handle.
Mechanical things don't fail as evenly as digital things.
There are a lot of simultaneous failure modes. HP had one in SMART firmware that killed disks at some fixed number of hours.
There are also temperature excursion events that can lead to arrays failing within 10s of hours of each other, not enough time to do a full rebuild of a single disk.
Seagate had really crappy 3/4 tb disks that loved to fail rapidly and nearly all at once.
If all your drives were Intel and part of the same array, they'd all fail at exactly the same time.
Reliable storage is hard.
And another poster os talking about SSDs?
Sigh
I make no assertion as to WD quality, but failures like you describe have been the bane of sysadmins quite some time.
I've worked at places where they always maintained a sizable inventory of new disks so that every time a new machine was provisioned, it would receive a RAID comprising of at least 3 disks bought from different batches at least a couple of months apart from each other, and as much as possible from different suppliers.
It happens. On my first job a batch of drives started failing daily, in sequential order by serial number. There was a bearing defect - the vendor ended up replacing about 400 drives, which meant many weekends of migration and restores.
Just because you were in a plane crash, doesn't mean it's common or frequent.
(With an apology to that guy who was hit with a nuclear weapon twice)
I don't deal with storage hardware much for some 6+ years, since I've switched to cloudy solutions. But at the beginning of my shared storage building journey, I considered arrays built from different disks somewhat "lame" or "unprofessional". I changed this viewpoint 180° after seeing multiple similar drives fail in a very short (hours, days) timespan for the second time.
I now have more trust for home-built ZFS setup using different disks than a dedicated storage appliance using enterprise-grade disks with almost following serial numbers.
https://docs.oracle.com/cd/E18752_01/html/819-5461/gbbwa.htm...
Given mine were in a ZFS mirror, they would have had almost identical wear so I suppose we could say they that they're predictable (though mine are the CMR drives which you don't get in the 4T these days and SMR drives are bad for ZFS).
If I was to replace them with the same drives I'd have a rough idea when to replace them before they start failing but I'll be buying larger capacities these days so those numbers are probably out the window.
For the story, I have a NAS with an AMD B550 motherboard and some Asmedia SATA controllers with a bunch of WD HDDs, using ZFS. At some point, they also started to throw weird read errors when under heavy load, but would work just fine the rest of the time. Turned out it was caused by me moving the SATA card from one of the PCIe slots wired directly to the CPU, to a PCIe slot wired to the B550 chipset.
I did replace them with WD, though, I think. Oops.
Is this multiple drives purchased around the time failing at the same time? Or is it a bad cable or controller?
If all of them went at the same time I'd try swapping to a different controller before I suspected the disks.
For the record, they have 43k power-on-hours, though nothing has been written to them much for the last ~3k hours as I had feeling something was off before they started reporting errors and migrated the data away.