A "clean" setup would include those, as well as either a messaging system or a regular checkup on how your FS is doing.
A "clean" setup would include those, as well as either a messaging system or a regular checkup on how your FS is doing.
They should really move to RAID-Z3 or 8 disks group. With RAID-Z3, we are looking at at least 3 disk fail at the same time with 15 disks around 0.04% chance. With 8 disks, we are looking at 0.2% chance.
So sure, the failures are not equally distributed and you've done well with your 184TB pool. Are you tracking the device errors, or only those that are visible to the OS? But multiple disk failures are not particularly uncommon. My experience in this space was two 16 disk servers that I set up 6 5 disk RAID5s (with one global spare per server). Within one month I had 11 of 32 disks die, and barely managed not to lose any user files, and this was not during the 1st month in production.
Scary. I've since moved to pairs of servers cross connected to pairs of 60 disk chassis (16 x 12 gbit connections per chassis) with ten 11-disk RAIDz3, 10 global spares, and 6x3.2TB of NVMe cache per server.
But the article is using the spec sheet URE rate which I'd assume looks only at the drive and doesn't take into account problems with the computer around the drive or the EOL time after the drive warranty has expired, I'd assume it was the "baseline" error rate.
> Are you tracking the device errors, or only those that are visible to the OS?
If we're talking URE like the article, that's data-loss on a disk, and the OS would always figure it out, since it would cause a ZFS checksum failure on scrub.
In this case it's not my data and not my money, so my preference is 6-drive RAIDZ2 vdevs. We've only had one disk with errors (and that one was migrated from a PC where Windows never reported any errors... of course...). The oldest 2 disks (3.5 years power-on time) have single-digit reallocated sectors in SMART so those are on course to be replaced.
I'm just curious since the argument in the article doesn't add up in my eyes.
> Within one month I had 11 of 32 disks die, and barely managed not to lose any user files, and this was not during the 1st month in production
Wow, that is some terrible luck!
Yes, they should have been scrubbing their pools, but I don't think this was bit rot. Millions of data errors is not what bit rot looks like. This looks exactly like bad hardware.
It is just bad hardware, bad setup and bad everything all cramped together.
But, agreed, whatever it was, if they were paying attention, they could have diagnosed and remediated well before they had any data loss.