What SMART Stats Tell Us About Hard Drives
backblaze.com
backblaze.com
#!/bin/bash
smartctl -a /dev/sda > /root/smartStates
grep Reallocated_Sector_Ct /root/smartStates > /root/stats
grep Current_Pending_Sector /root/smartStates >> /root/stats
grep Offline_Uncorrectable /root/smartStates >> /root/stats
grep UDMA_CRC_Error_Count /root/smartStates >> /root/stats
touch /root/statsOld
cmp /root/stats /root/statsOld
result=$?
if [[ $result -ne "1" && $result -ne "0" ]]
then
echo "Something went wrong"
exit -1
fi
if [[ $result -eq "1" ]]
then
echo "Files are different\n"
cat /root/stats
fi
mv /root/stats /root/statsOld
rm /root/smartStatesFor example, reallocated sectors are not alerted on by default, so we added '-R 5!' to our smartd config. The full config we have is:
DEVICESCAN -a -s (L/../../6/01) -l selftest -l error -m <email> -M daily -M test -R 5!
I feel like there were enough tools for users to simply monitor SMART stats, or awareness of how it works. Even in this article, it seems like a lot of analysis to see if reported flags are significant.
Add a "set -e" to catch errors. Say if the disk can't be read or file can't be written.
Why reuse the same temp file? Make a new one with mktemp and auto clean it via an exit trap. As it's written this isn't concurrently safe.
Exiting -1 on error? Don't use negatives.
Wrap it all in a main() function and use locals instead of global vars.
smartctl -a /dev/sda |
grep '\(Reallocated_Sector_Ct\|Current_Pending_Sector\|Offline_Uncorrectable\|UDMA_CRC_Error_Count\)' > /root/stats awk '/Reallocated_Sector_Ct|Current_Pending_Sector|Offline_Uncorrectable|UDMA_CRC_Error_Count/ {print $2,$10}'I don't see how following best practices for bash scripting (or really shell scripting in general) can be compared to J2EE standards.
Seek quality in all your scripting so it becomes the norm. Otherwise you'll end up having crap like that in something mission critical.
/var/log/syslog:Oct 6 08:14:10 hostname smartd[573]: Device: /dev/sda [SAT], SMART Usage Attribute: 190 Airflow_Temperature_Cel changed from 61 to 62
/var/log/syslog:Oct 6 09:44:10 hostname smartd[573]: Device: /dev/sda [SAT], SMART Usage Attribute: 190 Airflow_Temperature_Cel changed from 62 to 61
/var/log/syslog:Oct 6 10:14:10 hostname smartd[573]: Device: /dev/sda [SAT], SMART Usage Attribute: 190 Airflow_Temperature_Cel changed from 61 to 60
/var/log/syslog:Oct 6 18:44:10 hostname smartd[573]: Device: /dev/sda [SAT], SMART Usage Attribute: 190 Airflow_Temperature_Cel changed from 60 to 59My hard disks are fine, so right now the only stat changing is that, so that's what `rgrep smartd /var/log` turned up, so that's what I posted.
Did you consider looking at `man smartd.conf`? Did you consider that reporting and logging the temperature is entirely optional? Did you consider that you can configure it to log and report and email warnings for exactly the stats that you want? Did you consider that, for example, the stock Debian/Ubuntu package will pop up a window in your X session when stats which actually indicate potential failure change, like sector reallocations, read errors, etc?
Nah, just jump to the conclusion that all it does is spam logs with temperature readings, and downvote away!
Take Backblaze's 0.01% for a group of four failures. That's replacing an extra 100 drives per million, at random, and only getting the benefit of correctly predicting failures 10.4% of the time.
This is great data to have.
PS: Rememebr all RAM gets bit flit errors over time. Which is one of the reasons rebooting is often so useful, but also means one off errors are often meaningless.
As a personal user it's a non issue but scale things to ~70k devices and you get ~1.7 million device hours per day.
I dug and found https://www.fiala.me/pubs/papers/sc12-redmpi.pdf the title of which is "Detection and Correction of Silent Data Corruption for Large-Scale High-Performance Computing."
They state that for a cray there was a double bit flip about 1x/day for 75k modules. To (probably incorrectly) extrapolate, if you have a server with 16 modules that would be equivalent to a single double failure about once every 13 years.
If you want to have a quick, but in-depth look at your drives, it'll give you lots of data, including the SMART table interpreted in a vendor-specific way. It also understands some RAID setups, and more support for this is upcoming. Windows only, at the moment.
To explain a bit of a context - SMART data comprises a set of attributes and each attribute has a value, a threshold and a raw value. Values are opaque 8-bit somethings that are only meant to be compared to thresholds. When they fall under then, then it may indicate a problem. They aren't really interesting. What's interesting is the "raw" values, but as the name implies, they are vendor-specific and require decoding. Some vendors publish the specs, but most don't. Specs that are published are often incomplete or plain wrong. So there's a LOT of reverse engineering and guesswork involved, which makes writing a SMART tool both frustrating and interesting at the same time. But if you need just the "dying / healthy" indicator, it's a very easy thing to extract from a drive.
Not the SMART part, but how you talk to the drives and controllers and how storage is generally sliced into partitions, volumes, etc. Windows has a fairly comprehensive version of Software RAID, but in true Microsoft fashion they do things ass-backwards in more than one place. For example, striped volumes (RAID 0) will use only a part of a partition for each stripe, but to learn that you'd have to talk to Virtual Disk Service rather than regular Disk/Volume management API. This is, basically, as unportable as it gets.
Such nuisance for what might be a good read.
Secondly, I noticed that this is a trend right now; where you get to a page and after a few seconds a dialog just thrown into your face with little disregard to you (the reader) is trying to concentrate to read the content. To me that is rude, you don't go to a bookstore while reading the table of content a salesman grab that book from you and tell you "would you like me to take your email address so that we can notify you when we have new books available?" without wondering what kind of establishment that allow this kind of behavior.
Third, I got the link from HN it was easier for me to go back to this tab, login, hit reply than registering a disqus account and then enter a comment there.
With that said, I dont want to blog about this on Medium or whatever, I dont need clicks by moaning about every little things, this is my way of protesting on what I perceive is happening right now and that's why I start with "small nitpicking".
Without people leaving the sites and complaining on places like HN, web designers will have no feedback that it's such a stupid idea.
But I don't know that until I've already given up the goods, do I? This approach to life, the universe, and the Internet simply doesn't scale.
I hate ads. I block all of them. But I will help your crowd-funding or Pateron or buy some swag to help you promote your thing.
@blackblaze I'm pretty sure you can automatize a large portion of your investigation that way.
Second, timeouts and uncorrectable errors are generally being reported to the controller as part of normal operation. So having SMART tracking them is just a bonus. Either of those two conditions is usually sufficient to kick a drive out of a functional RAID array because those are data loss events. Most drives have layers and layers of ECC, so in order to get an uncorrectable error a lot of bits need to be flipped in the target sector. For that to happen it likely indicates there is something mechanical going on which is likely to affect adjacent tracks/sectors. Of course if you never scrub your drives its possible bitrot accumulates on a perfectly functional device until sectors aren't recoverable.
In my previous life I found it much more interesting to track the rate of soft error counts during scrub operations. Particularly, in larger arrays because sometimes a drive would start getting slower (which is frequently caused by read retries in the drive itself or problems tracking the embedded servo/etc) and the correctable error counts would start to steadily rise followed by actual timeouts/uncorrectable errors. Of course these days, it seems most drives won't show the correctable error counts because it would freak people out. Instead you have to infer it from seek errors and relocated sector counts. Although, it might now be considered a SAS/SATA differentiator. SCSI has standardized log pages with more detailed information. (random google hit http://www.seagate.com/staticfiles/support/disc/manuals/scsi... page 238) Note the errors are categorized as corrected without delay, with substantial delay, and corrected on a retry. By comparison the SMART data isn't particularly "smart".
http://static.googleusercontent.com/media/research.google.co...
I suspect this is because value, worst, threshold columns are kind of confusing to understand.
SSDs though, they just disappear from the bus when they fail; so I haven't been able to look at a dead one and see what looks like a useful predictor. I have seen some ssds reallocating a big block, which kills performance while its going on...
This isn't always true, and actually shouldn't ever be true - it's a particular failure mode you're seeing, and while it appears to be one common across a number of SSD controllers, it's still a pretty sorry fact that it happens.
All SSDs (at least all not-complete-rubbish ones) report some kind of flash/media wearout indicator via SMART, which isn't necessarily an imminent failure indicator (SSDs will generally continue to work long past the technical wearout point), but is a very strong indicator that you should replace it soon and should probably buy a better one next time.
SSDs do suffer from sector reallocations in the normal way, and the same kind of metric monitoring can be done. It's pretty vendor-specific as to what SMART attributes they report, but attributes like available reserved space, total flash writes, flash erase and flash write failure counts and so on are pretty common.
SSDs on the other hand ...
I use SSDs for caching (ZFS read cache and mirrored SLOGs) and I use them for mirrored boot devices in modern, production systems that should have a fast OS device.
But if I want a system to run forever ... if I am optimizing for longevity ... I use compact flash, even in 2016.
(yes, of course I set them to be read-only and disable swap)
Would you enable an option for smartmontools that sent all your drive SMART data to a cloud hosted db (with as little identifying information as possible) and tell it when you had a drive fail - in return for that same service alerting you with "best estimates" of your risks of drive failure?
This doesn't gel with the prior table, which shows that 4.8% of operation drives have non-zero SMART188 alone.
I could easily see one set based on a different time, or another that missed a category of drives (EG one also counts drives from testing / non-production units).
Perfect predictability would obviously be beneficial, in that you could get by without any redundancy. But even imperfect predicability can help you reduce the required number of drives for a set level of security.
Not perfectly, obviously, but this is about probabilities.
Say you want some level of data security (i. e. 99.9% over one year).
The formula for the risk of data loss is (r^n) where r is the failure rate of drives and n is the number of drives (assuming independence and using just mirroring).
But – if you have a test that can predict failure with some probability s the formula becomes (n*(1-s))^r. which is strictly smaller for any 0 < s <= 1, meaning you get a higher level of expected security (or possibly the same with fewer drives).
Since it's all probabilities, triple redundancy does not guarantee complete absence of data loss. On the other side of the spectrum, single replication might offer a better cost/safety ratio for some applications.
A failure that is known in advance is equal to no failure in these term and even when discrete, it will make the difference once in a while.
I think it's obvious. If the drive has errors, this may cause or cause need for a reboot.