This post is a very strange failure mode so it'd probably not be in that guide.
Iffy sectors get marked as Pending, Reallocated sectors are also marked. There's some smart stats for seek times too, I believe (but those weren't as clearly associated with failure as sector counts).
There is a danger that manufacturers will avoid acknowledging problems in smart, but so far, I haven't heard about that happening too often. There are certainly some drives that are dead or dying with fine smart values, but it seems rare, and there's often a not obvious dependency where a drive that is very unhealthy may not be able to record new smart values.
- Seek time gets very high (>80%). - Drive drops dead and won't respond
SMART doesn't detect either of these.
Setting an alert on seek times > 30ms has been by far the best predictor of drive failure. If the seek times go over ~400ms then other components start complaining about slow/missing I/O, but by that point you've already gone crazy because your array is unusably slow.
If a drive develops bad sectors SMART will report it, but that's rare enough that I don't bother watching for it. It's a non-event with ZFS because of the checksumming. These drives never have one random error; they start raising hundreds in the space of minutes, and that's the moment you replace the drive.
thing is, "iffy" depends on what firmware decides is iffy.. Maybe a firmware relocates a sector the very first time it fails to read, maybe it relocates after try 200. What we want is to be notified, when we request a sector, if reading it was 100% painless. Depending on your demands for reliability, maybe it's okay to re-read a sector a few times, and still not reallocate it, but maybe your filesystem would like to gather these stats from the disks anyway.
I can understand why the default mode of operation for disks involves the disk presenting a clean “you get a block or you don’t” interface to the OS over e.g. SATA; but I’m surprised that it’s what we’re stuck with 100% of the time. We have ECC memory (where the memory controller reports errors to the OS), but no ECC disks.
Maybe this is what Apple is trying to do with putting their T2 chip in charge of being a storage controller. That way, it doesn’t just get to do encryption things; it can also make highly abstract policy-level decisions on what to do about these sorts of error events. Too bad such an approach is only really tenable if you’re building your own storage system out of your own raw NAND; it’d be nice to be able to have an external low-level storage controller built into e.g. the RAID card of a RAID array of regular HDDs, that those HDDs would then slave themselves to, and which ran OS-uploadable firmware.
We deal with this by watching for these errors, printing to a log specifically for Icinga to watch for and alert on, and preemptively replace the disks. It would be nice if the other software (ZFS, SMART) would notice these in time to not become severe.
> In a typical case, if you don't cook pork well enough, you digest live tapeworm larvae. They've got these little hooks, they grab onto your bowel, they live, they grow up, they reproduce. Reproduce? There's only one lesion, and it's nowhere near her bowel. That's because this is not a typical case.
New professionals always expect things to be similar to what they studied. What they learn: the books and lectures often covered the "most common" situations, but all those edge cases and weird scenarios come up on a daily basis.
[1]: https://www.springfieldspringfield.co.uk/view_episode_script...
The gist is that in medicine there's the concept of diseases that are zebras and horses, rare and common, and the saying goes "when you hear hoof beats think horses not Zebras". Unlike medicine, in software most the easy problems have already been abstracted away or made trivial through our software stacks, so often what's left are Zebras. What used to be the exceedingly rare bugs now account for a large proportion of the odd behavior we see in very robust tech stacks, because the easy bugs have been hunted down and eliminated already.
Boris Beizer, Software Testing Techniques. Second edition. 1990
One obvious caveat to virtualisation is the overhead, but with time, you can learn how to use virtualisation with ease and re-provision an entire new OS (and virtual disk) if needed. In terms of how this relates to disk health: well the answer is you now have multiple healthy disks, and only rarely have to then deal with failing disks, which due to compartmentalization, the risk of a failing disk ruining your day is minimized and hopefully the failing disks have nothing too crucial on them. Throw in some cloud storage solutions for backups and this makes things even more bareable.
Aside: One thing I've done a few times is setup a queue and service to replicate data in an rdbms to elasticsearch or mongo as completely denormalized records for searching against. When I did that, it became very easy to store a .json.gz of these same records in s3/blob storage as a secondary backup.