Predicting Hard Drive Failure with Machine Learning
datto.engineering
datto.engineering
It’s a bit of a breath of fresh air to see a post like this where someone clearly puts the work in to try and solve a given problem with ML and is able to articulate the failure modes of the approach so well.
For example, in this post the engineer does not adequately handle class imbalance at all, and reaches for a hugely complicated model framework (XGBoost) without even trying something like bias-corrected logistic regression first.
I think this post is just an example of an infra engineer thinking “ML is easy” and when their quick attempt doesn’t work, then just bailing and saying “ML is overhyped and doesn’t work for my problem.”
Stuff like this really should be handled using survival analysis.
Too much valuable information is lost because people don't value negative results enough.
There are some quite natural questions you can't answer unless you can predict some stats on when failure will occur, and if you have any drift in the dataset, you might just find features correlated with age (and hence having more time to fail) that don't inherently increase risk.
I would recommend:
(a) using survival analysis
https://lifelines.readthedocs.io/en/latest/
https://xgboost.readthedocs.io/en/latest/tutorials/aft_survi...
or (b) flattening the data to be one observation per time window, and predicting failure over specific time windows.
This is a good talk also. https://dataorigami.net/blogs/napkin-folding/18308867-video-...
Although I'm not sure about this case, in general the approach in the article (tree-ensemble predicting binary classier) works well for many business problems!
Thermal abuse (running at a high temp for many hours), duty cycle, iop rate (high head acceleration versus head dwell) all impact the lifetime.
Early life media issues are predictive of failure and can be detected through stress testing with various seek ranges. It’s not as simple as a SMART metric, but did appear to have predictive value. The one experiment I ran discovered a set of drives which ultimately the manufacturer admitted to a since corrected process problem.
The above factors are different on each drive family and can’t normally be generalized.
I could search or visit one of the social pages linked to in the footer but it seem like an obvious link one would want to have.
[*] The underlying neural networks aren't deep at all.
The punchline here is: > Unfortunately, once they were scaled up to production levels of use, their tradeoffs turned out to be too much to bear. These models are good—especially given how imbalanced the data for this problem is—but with too many false positives from one and too little predictive power from the other, they’re not worth betting so much money on.
A practical next-step analysis would be some back-of-the-envelope economic math about the necessary detection threshold for detecting and replacing failing drives. Overall, it doesn't even make sense to even think about applying ML without understanding the cost/benefit of solving the problem.
Think about Backblaze -- the drives are packed in there tightly and it's super expensive to replace one, whereas there's basically zero cost to leaving one in place. Even if you could perfectly predict which drives would fail, there's no reason you'd act on that info.
Failure Trends in a Large Disk Drive Population, FAST '07
https://static.googleusercontent.com/media/research.google.c...
> Each attribute also has a feature named raw_value, but this is discarded due to inconsistent reporting standards between drive manufacturers.
Sure, raw_value is inconsistent between manufacturers and sometimes models, but it's the most real data.
Edit to add: reread the Google survey, which says more or less bad sector counts were indicative of failure in about half the cases; in 2007.
One thing to note is drives made now may perform differently than those in the 2007 survey. Different firmware, different recording, better soldering (200x is full of poor soldering because of RoHS mandating lead free solder in most applications, and lack of experience with lead free solder leading to failing connections).
I found, in 2013ish, with a couple thousand WD enterprise branded drives that just looking at the sum of all the sector counts and thresholding that was enough to catch drives that would soon fail. If your fleet is large enough, and you have things setup so failed disks is not a disaster (which you should!), it's pretty easy to run the collection and figure out your thresholds. IIRC, we would replace at 100+ damaged sectors, or if growth was rapid. Some drives in service before collecting smart data did seem to be working fine at around 1000 damaged sectors though, so you could make a case for a higher threshold.
SSDs on the otherhand seemed to fail spectacularly (no longer visible from the OS), with no prefailure symptoms that I could see. Thankfully the failure rates were much lower than spinning disks, but a lot more of a pain.
If the stats that matter only update after a test, and noone runs those tests, then is SMART the failure or is the lack of caring the failure.
>Only one drive manufacturer had enough failed drives in the dataset to produce a large enough subset, which we’ll refer to as Manufacturer A
Seagate, we will refer to this manufacturer as SEAGATE.
I wonder if there are any existing examples of such practices. Supermarket chains sell products near their expiration date at steep discounts, of course the difference is you know what you are buying and can still eat a 20 days until expiry 50% price Nutella jar.
If your system is designed for hot swap of failed disks, but 1% of the time, you need to reboot the system for the new disk to be detected; predictive replacement lets you move traffic off the system at a convenient time to do the swap (when traffic is low, and operations staff isn't super busy). Predictive replacement can be deferred more easily, because the system is still working; maybe you can batch it with some other maintenance.
A surprise failure needs to be dealt with in a more timely fashion.
In my experience (which I didn't write a blog post for), monitoring the right SMART values and doing predictive replacement eliminated most of the surprise failures for spinning drives. SSDs had less failures, but I wasn't able to find any predictive indicators; they generally just disappeared from the OS perspective.
Also we haven't even discussed the online disk repair phenomenon, in which you take a "failed" disk, perform whatever lowest-level format routine the manufacturer has furnished, and get another 5 years of service out of it. This is done without ever touching it.
I've seen some drives that were apparently working fine with about 1000 reallocated sectors; and the smart firmware usually indicates a maximum of quite a few more than that, so if you replace at 100 problem sectors, I would consider that 'predictive' replacement. It's debatable of course; something that predicted failure based on changes in data transfer (that presumably precede sectors being flagged), would be a bigger predictive step, but I never got that far in modeling; using flagged sector counts was enough to turn unscheduled drive failures into scheduled drive replacements for me.
What I wonder than, is if sensitive isolated microphones could be tried for this purpose? We already know that sounds (ie yelling at the drive) can vibrate the platter enough to cause performance degradation. If there were internal mics in each HDD recording sound as the HDD spins and correlate that to HDD activity, could that be correlated with failure rate?
If you have good operations, replace when convenient after it hits 100. If you have poor operations (like in home use, with no backups and only ocassional SMART checks) replace if it hits 10 and try to run a full SMART surface scan before using a new drive.
For SSDs, good luck, I haven't seen prefailure indicators.
Current_Pending_Sector + Reallocated_Sector_Ct + Offline_Uncorrectable + ???
What I'd be really curious to try is run IO benchmarks on the drive and see if there are performance issues that indicate a drive is failing.
This was written as an after-the-fact look back after a year long project and the target testing set did change a little along the way. The big change was splitting by manufacturer (and eventually by model family), which changed the number of applicable samples to compare performance on.
I cut some stuff from the article for length, so I totally get how that could have been unclear.
Some hard drives have a vibration sensor(s) in them - i wonder if a daemon could sample throughout the day and dump a summarisation of the data when the daily SMART data dump is being generated.
The drives report temperature via SMART, i wonder if anything meaningful can be extracted from the temperature differential between the HDD’s sensor and a chassis temp sensor?
I wonder if more discrimination by manufacturing detail (date of mfr, which factory etc) could help?
I was going to suggest drive supply voltages but those are hard to compare across chassis - calibration of voltage measuring circuitry in PSUs isn’t too hot for the millivolt differences you’d want to measure - the differences would be so small to be lost in the inaccuracies between boards. Also many PSUs lack a method to query their measurements.
Fascinating study and excellent writeup. I have zero ML background but i really enjoyed this.
In the bio blurb they self-describe as an infra engineer who also enjoys data science. In some sense I really don’t like to see that. The quality of the data science / ML work in this is actually quite bad, but people use these blog posts as resume padders to try to jump into ML jobs without ever having any real experience or training.
I think it’s a bad thing because it devalues the importance of real statistical computing skills, which are very hard to develop through many years of education and experience - absolutely not the sort of thing you can get by dabbling in some Python packages on the weekend to do a little project like this.
The amount of waste I see from companies trying to avoid paying higher wages and avoid team structures that facilitate productivity of statistics experts is staggering - with all kinds of hacked up scripts and notebooks stitched together without proper backing statistical understanding, making ML engineers manage their own devops, and just ignoring base statistical questions.
For this drive problem for example, I expect to see a progression from simple models, each of which should address class imbalance as a first order concern. I expect to see how Bayesian modeling can help and how simple lifetime survivorship models can help. I expect to see a lot of feature engineering.
Instead I see an infra engineer playing around with data and trying one off the shelf framework, then claiming the whole premise can’t work in production.
This article is the example of wasting time and money. It’s amazing to me the way anti-machine-learning sentiment causes people to do a complete 180 from common sense just to avoid actually investing minimal levels of resource or effort to study and understand problems where statistics can help.
A big observation in my career is that statistics makes non-statisticians go crazy in the head. People panic that statistics will be used to usurp their authority and then try to steamroll statistics with politics about what is or isn’t over-hyped and what would or wouldn’t be a “waste” of time or resources to try out.
I'd have dropped Manufacturer A from the supplier list and used the A only model for the remainder of their drives. Then I'd have another go with SmarterCTL for "not A".
But if the drives are part of redundant array, it would be almost always cheapest to let them fail.. and if large part of failures are asymptomatic, you need the array anyway for critical stuff, so I suppose it's a useless exercise.
For example, perhaps the read errors number increases about once a week. But in the week before a failure, it increases once a day. Or in the day before a failure, it increases once an hour.
OS-level metrics could also be predictive: write latency, i/o queue length, communication errors, and checksum errors in read data.
You can filter out some of the bad drives ahead of time, but you will still get blindsided much of the time.
As for the rest of the counter, raw read error rate, relocated sector count and friends are also good indications that your drive might not be in the best shape. They don’t say anything about when it will fail though. I’ve had a drive rack up 20 relocated sectors in a couple of weeks, only to stay at 20 for 4 years after that.
And load cycle count is only good for telling if you’re within manufacturer parameters. I frequently see old drives with Load Cycle Count 10-25 times the rated value, and they work just “fine”.
1. Extreme imbalance and survivorship bias. The dataset is unwieldy because of the imbalance, no matter the algorithm; and too small if you balance it.
2. I would take a daily/weekly/monthly diff between different values and time to failure, and obviously only from drives that ACTUALLY failed. All drives fail, but data from healhty drives is useless as it lacks the target value. In practice you want to predict how close they are to failing and act preemptively, tagging them as OK and KO completely misses the target (hehe) IMO.
3. Drive maker and model should be vectorized and fed into the model. Mandatory. Different manufacturers and models show different SMART behaviors, and while some stats are red flags in certain drives, they mean nothing in others.
4. This is more of a personal preference, but I like random trees to test whether the dataset is useful. They're not perfect but they give reasonably good results for most tasks, and tend not to overfit, so you can use it as overfitting benchmark vs other algorithms. If they don't either the dataset is shit or it needs some transformation, or the model needs to be more complex (but that's easy to see). Obviously this applies to classifiers and regressions, and not to many other ML tasks.
What is the right solution at consumer level though? I currently run smartmontools at regular intervals, compare the TBW with manufacturer's TBW and send a MQTT message using a script.