Hard Drive Reliability Update – Sep 2014
backblaze.com
backblaze.com
Same for the 7200.11: http://knowledge.seagate.com/articles/en_US/FAQ/207951en
Isn't using very similar drives a problem because their failure rates are not statistically independent, so there is an high probability that they will all fail at the same time?
There are some statistical properties I recall that will make mechanical failures spread somewhat nicely, maybe someone can elaborate if they have a background in stats.
I would guess in your scenario something like that happened. Or perhaps trying to migrate tons of data to a new disk caused an issue. Just seems unlikely otherwise.
I guess you could calculate confidence intervals at quarterly intervals, and so the error bars would get larger as age increases and 'n' decreases.
How would you calculate the CI for failure rate? It's not binomial or poisson, since failure rate goes to 1 over time...
A little searching turns up http://rmod.ee.duke.edu/statistics.htm which I'm sure completely explains how to do this... (rolls eyes). I hate that this is how statistics is commonly taught. Knowing which distribution to use and applying it correctly can actually be intuitive if taught properly. It doesn't always need to be an exercise in alphabet soup / deriving from base principles.
This is similar to studies in medical research, e.g., how long do subjects live after an experimental cancer treatment. What I've seen used is the Kaplan-Meier non-parametric method (Kaplan-Meier plot) for the purpose.
Maybe I'm just used to seeing these plots, but I find the method conveys information very effectively. More info can be found here: https://en.wikipedia.org/wiki/Kaplan%E2%80%93Meier_estimator
I agree that presenting the SSD life data in that way would be a good idea.
EDIT: More fine grained data is probably needed, granted.
Grad-level version: drive lifetime is exponential with an inhomogeneous (increasing, presumably) rate function lambda(t). Inferring lambda(t) is difficult without additional assumptions on the functional form. But potentially do-able.
Real-life version: None of the classical distribution fit that well. (https://www.usenix.org/legacy/event/fast07/tech/schroeder/sc...)
(This is survival analysis: http://en.wikipedia.org/wiki/Survival_analysis)
They showed some of this data in an earlier post https://www.backblaze.com/blog/how-long-do-disk-drives-last/
The "Drives have 3 distinct failure rates" graph is the most interesting, as it shows the result of the expected "bathtub curve" on the cumulative failure rate.
https://www.backblaze.com/blog/what-hard-drive-should-i-buy/
The second takeaway is: Check the last Backblaze report right before you buy new drives.
/HGST employee
It is so bad even running HD video off of some 2014 HGST laptop hard drives causes distortion. There's also no way to increase the ADM value indefinitely that I am aware of (it reverts after every sleep).
For those not "in the know" the hardware is the same, but desktop firmware drives will sit there for 10 seconds or whatever it is beating the drive when there's a read (or write) fail on the assumption that if your machine only has one drive you're better off trying as hard as possible to keep retrying until it works, and possibly the slowness will motivate them to replace (god forbid an end user have backups lol)
Enterprise firmware, when it has a soft fail, just croaks as fast as possible. That lets the raid array hurry up and do its thing, or maybe even higher level replication do its thing.
(edited to add the old startup adage of "fail quickly". Thats what enterprise drives do to keep overall array latency low, which is counter productive for consumer non-array drives)
Aside from the firmware load the prices are different because usually enterprise has better guarantee and better service and unlike consumer drives which statistically are never replaced under guarantee so you can claim anything on paper for marketing purposes it won't cost anything, enterprise drives WILL get replaced and there will be a papertrail etc. So the guarantee for an enterprise drive actually costs something.
Sometimes the firmware has some other subtle differences like how it handles recalibrates and scrubs (consumer home drives are like "too bad you get to wait on my schedule" and again, enterprise will go to some effort to eliminate array latency)
My guess is the article is subtle astroturf by the winning drive mfgr?
I imagine someone in the drive / enterprise storage business would know better, but many of us might not have known that.
Oh I should preface my remarks with "in family, same technology level" etc.
I'm talking about the same package with the same storage where different firmwares are loaded. Thats the business model the OP is in, he makes posts about stripping consumer drives out of external cases and I don't want to speak for him but he radiates the impression of their whole secret sauce being reliability at the system level not the hardware level.
It is true in my experience that "exotic hardware" marketed as enterprise grade is different physical hardware and is going to be a different level of reliability than consumer hardware.
So the difference between an office NAS SATA 1 TB enterprise and consumer grade is a firmware load, and I think thats what the article is talking about. There is a big difference between a 15K SAS enterprise server drive and a 7200 SATA consumer drive at the hardware level, but that's not the article author's "thing".
They're never replaced because regular consumers statistically don't ask for replacement and simply eat the loss?
At a business that is probably some MBA metric and has been budgeted for and is salaried anyway, and you need to prove I'm not just taking drives home to put in my basement server, so they have to be destroyed or sent back, and there may or may not be PCI / CPNI type concerns, so yeah, they get sent back.
drives with longer warranties generally cost more in the consumer segment, and data like this could help people identify the "sweet spot" balancing cost with risk of failure, warranty length, and depreciation curves for storage. my "ancedata" suggests somewhere between 2 and 3 year warranties seems about right for consumer drives these days...
There is another paradox as well. Some people won't ask for replacement because assuming you have to send in the bricked drive there is the chance that someone might get at your data somehow.
What about that? (It's why I would never send in a drive that has failed.) [1]
[1] My assumption is that I would have to send in the bad drive (and there is no way to reformat or for me to easily destroy what might be on there). Anyone have experience with what happens here?
So for example you might want to encrypt something super sensitive (which I do) but decide to not encrypt something less sensitive (say photos or perhaps a wiki with notes or letters to your grandma or wife or sig other).
Point being that if the drive isn't encrypted you might be able to get at some of that data. (If you need to). If you've encrypted the drive then can you still do that?
I strongly recommend Tarsnap[1] for that. All your data is encrypted before it leaves your machine, it is run by our very own 'cperciva, and the key used for encryption (which you need to store securely somewhere, like your parents' house or a bank) can even be printed so hard to destroy accidentally.
The key itself can be encrypted, too. You can use the same password you use to encrypt the drive and now you safely and securely store all your data in such a way that only you can ever access it by remembering a single, longer phrase.
(Although to be fair, another backup would probably be a good idea if the data is really important. Maybe another encrypted hard drive kept at work.)
I'd like to see on sites like that some kind of continuity plan.
Unless you are suggestion that this is just another "redundant array of backups that you have to assume can fail".
But even in that case it would be a good idea if cperciva had something posted on the site which showed there was someone else who had access and kept on top of the system.
Because now I have a problem: should I believe guys who have been running 38 petabytes of storage for several years now and regularly present the data they gathered, or should I believe you?
I am open to the idea I'm misinterpreting how he presents enterprise vs consumer firmware loads. I assume its open knowledge that their secret sauce is using consumer hardware so assumptions about 15K fibre channel isn't relevant.
Can you respond to this directly, please?
o Enterprise and Consumer Drives are the Same Hardware.
o Enterprise Firmware causes the drive to fail fast on physical errors
rather than endlessly retrying.
o Enterprise Drives are more likely to be RMA'd?
Do you have any citations, evidence, reports, articles, white papers, research - anything (beyond random anecdotes, or miscellaneous blog entries) to back up these claims?http://features.techworld.com/storage/1019/what-is-time-limi...
TLER is exactly what you want in a RAID. You have another copy of the data in question on the array, why fight for it on a disk when the data may be corrupt anyway.
What I'm interesting in hearing, (honestly - I not doubting right now, just interested in being educated) - is if anyone authoritative has described "Enterprise" drives as having these behaviors.
Except that also only works "most of the time".
Anyone who works at scale with drives and (RAID-)Controllers knows that even "enterprise" drives can and do take entire controllers down.
It's not uncommon to lose a full set of daisy chained JBODs to a single disk acting funny.
Consequently, and since storage clusters have to be redundant at the node-level anyway, it makes a lot of sense to skip the markup for enterprise firmwares and instead design for quick node failure detection and ejection (short timeouts).
Disclaimer: HGST employee.
I'll leave this non-HGST (Intel) link (PDF) for reference:
http://download.intel.com/support/motherboards/server/sb/ent...
The concrete features I understand "enterprise" drives as having are:
1) firmware is built with RAID in mind. This might sound weird, but consumer drives are more likely to have problems with RAID, just because they aren't designed for it. See, for example, the first WD Green drives crashing RAID clusters because some timeout was too high.
2) Longer warranty period
3) This may not be true anymore, but I believe they typically feature larger cache
4) ECC
[1]: http://www.pantz.org/hardware/disks/what_makes_a_hard_drive_...
As for the RAID point, you'd very much want to find real data for that. I've heard similar folklore but have also heard plenty of lost data stories from enterprise disks in enterprise RAID. If I were building out a large storage farm, I'd strongly consider forgoing hardware RAID altogether and using something like ZFS so it'd be observable.
If you've ever worked with enterprise RAID, that should not surprise you! RAID is powerful, but also so easy to botch...
Dell, Google, and Amazon can never write reports like this because the vendor relationship is important. Because these guys have no relationship and are buying consumer disks, the world finally gets brand level reliability reports. Kudos to Backblaze.
Dell maybe, but Google and Amazon consider what goes on in their data centers to be part of their secret sauce. It doesn't have to do with vendor relationships, it has to do with a competitive edge. Google and Amazon will drop a vendor without a second thought if it's even slightly more advantageous.
The worst part of it is: my own HDs are all Seagates and I use them longer then that...
As others have said though, these things tend to go in cycles. One manu. gets a bunch of bum drives, people react and they start cracking the whip on QC. I think it's also a product of what we're asking of our drives today: SSDs are eating away at the low end so their only recourse is to pack more and more data into smaller and smaller spaces leading to a greater reliance on error correction.
I will say though, the BackBlaze report led me to try out HGST drives for the first time and I've been overwhelmingly satisfied with them.
That said, will that trend continue? I guess we'll see when the 2015 report comes out.
My experience has been that any rule about "brand X is the best" or "brand Y is the worst" only lasts for a short time--each hard drive technology "generation" seems to reshuffle the reliability matrix. Whoever was on top in 2009 has almost no bearing on who's on top in 2014, or who will be in 2019.
I got into building PCs on the side in about '96 or so and did break/fix support through college - it always seemed (I didn't keep records - I can't quantify) that Seagates accounted for more than their share of the "it won't start and is making a clicking noise" cases. So much so, that I have recommended against Seagates and towards WDs for as long as I can remember (15+yrs).
I've only had a small amount of direct experience with Hitachi drives, so I can't speak to them too much.. Though I do have a Hitachi in an enclosure that's been kicking around since about 2004, so I guess that says something...
†They apparently dropped this to 3 years for the first 6 months of 2012 at least for the big drives who's technology base is consumer ... I suspect that change was not well received.
Unfortunately, you are generally correct that we've seen a race to the bottom.
But when they released the 7200.11 versions....ALL 7200.11 models would spin down after idle for X minutes and then start clicking...the data was still intact, but for the drive to work again you had to pull the power.
I unfortunately built an eight drive RAID5 with these drives before the issue was know, which made it very difficult for me to diagnose the issue. (All drives seemed perfectly fine when powered and working for the first 5-10 minutes or whenever they were in use, but as soon as a single drive of the RAID idled, the RAID5 acted like a drive was bad).
I still have a box with eight 1TB Seagate 7200.11 drives that I never updated. I have never bought another Seagate drive for RAID since, always WD or Hitachi.
However, they are cheap, and they do honour their warranties. Would just be nice if they didn't have to quite so much.
https://www.backblaze.com/blog/why-now-is-the-time-for-backb...
At these rates, why not use S3? What am I missing?
It is from our original blog post about our storage pods (we assemble 45 hard drives in some sheet metal with off the shelf parts):
https://www.backblaze.com/blog/petabytes-on-a-budget-how-to-...
We've now written over a petabyte, and only half of the SSDs remain. Three drives failed at different points—and in different ways—before reaching the 1PB milestone[1]
http://techreport.com/review/26523/the-ssd-endurance-experim...
http://us.hardware.info/reviews/4178/10/hardwareinfo-tests-l...
Hardware.info did a "Write until it dies" test for a couple of 250GB Samsung SSDs last year. They found that the drive consistently exceeds the 1000 writes per cell spec.
I recently examined a set of Crucial m4s which were at 130% of usage. There was no lost data, but write bandwidth was hilariously bad (around 10-20MB/s).
The SSDs helpfully keep track of how many bytes you've written and report that in the SMART info. For example, on my Windows dev system, HDD Guardian reports that I've used up 42% of my SSDs endurance (it's a Crucial m4 512GB). So by "usage", I mean the percentage of the endurance which has been burned.
[1] Each block of flash memory must be erased before it can be written. Each time you erase it, it "uses it up" a bit and will be harder to erase next time. So each time, the SSD controller is forced to erase it with just a bit more voltage. As the blocks become harder to erase, it takes more time to erase them and the SSD write bandwidth decreases.
Specs and data sheets
cox proportional hazards models and KM survival curves are the big kahuna with a data set like this. basically, my impression is that you'd want to pretend that you're doing a cohort study and essentially analyze it like a big clinical trial.
and re: graphing, since you have sub-samples of the different drive models and relatively different but still small numbers of models for each brand, the big take-away graph comparing manufacturers leaves out a lot of useful nuances that could be pulled out of the table, e.g. that most of the hitachi drives are newer and you have fewer of them. it also doesn't portray how consistent failure rates are across models produced by the same brand might be, and it looks like there's a significant range of failure rates across different seagate models. so even just seeing IQRs of pooled failure rates for each manufacturer in box plots might be eye opening..
immortal time bias may be something to consider here as well...when you have some sample groups that mostly include newer individual drives that you have not yet had for a year on average, subtle differences in how the failure event is described can make a big difference in the conclusions you draw...especially in terms of uncertainty. if i have 10,000 hitachi drives and 5 of them fail in the first 6 months, the robustness of conclusions i can draw from those data are different in some important ways from similar insights drawn from a sample of 1,000 hitachi drives i've used for 5 years.
it's also not clear to me how you've dealt with replacement drives. based on what i gather from the post (and i could be totally wrong), if you have a bunch of one model of a drive failing and then get replacements for them... some might argue that refurbed drives should be analyzed almost like a separate drive model, since they often have physical differences compared to those bought via retail channels.
i'd be quite curious to dig into these data a bit further if you're willing to post the original data set...
thanks again for posting information that's quite useful and interesting
Off-topic, but It's a shame that BackBlaze isn't available in some countries, I'd love to use it. What would be the best alternative to it, Tarsnap?
[1] https://www.backblaze.com/blog/what-hard-drive-should-i-buy/
I've been procrastinating about getting off-site backup. This post on HN reminded me that I've been meaning to get an account going with your company for a while. I just signed up and will test on my machine before deploying to other machines in my business. Thank you.
HGST and Wester Digital are the same company, but it seems they have separate product lines? It's confusing.
The 3.5" desktop and server Hitachi drives are all manufactured by Toshiba. WD wanted to buy Hitachi's drive division so there would be a WD/Seagate drive duopoly, European regulators said they had to sell Hitachi's 3.5" to someone else before they would approve the merger. Toshiba stepped up and got it.
http://www.anandtech.com/show/5635/western-digital-to-sell-h...
So yes it is somewhat confusing. You buy one of these drives and it will say "HGST a Western Digital company" on the box, when in reality it's all made by Toshiba, probably in an old Hitachi factory.
First, Toshiba had their own 3.5" drives long before the merger with Hitachi. For example, Toshiba MK2002TSKB 2TB 3.5" hard drive was on sale since 2011.
Toshiba's own design and factories wouldn't magically disappear after acquiring Hitachi's assets. So, after the merger, Toshiba sells some ex-Hitachi drives, for example toshiba DT01ACA300 3TB is obviously a relabeled Hitachi drive. But I believe that Toshiba MD03ACA300 3TB and Toshiba MD04ACA300 3TB drives are based on Toshiba's own design because they don't look like Hitachi or DT01ACA300. I suppose reliability and performance will be different between these three Toshiba 3TB drives. And it would be interesting to get some info on this.
Also, HGST is wholly owned by Western Digital. Despite regulators requirement, they didn't sell all 3.5" assets to Toshiba, only some of them. Most of the good stuff (that is, everything currently sold under HGST brand) still went to WD. So, even ex-Hitachi drives sold under Toshiba brand (by Toshiba) and under HGST brand (by WD) might have different reliability.
[No lost data, I do daily backups.]
I do however have a Seagate drive laying around somewhere that has almost 10 years on it and it still functions flawlessly. But it is admittedly smaller given its age and that may contribute to lifespan. Either that or Seagate has slipped in the last decade.
Hard drive failures tend to happen more for drives that are power cycled a lot, and for drives that undergo big swings in temperature (even when the temps are all within the rated temp range).
I'm just going to keep using Seagate until my anecdata refutes the reality I live in.
Do you burn in new drives before using? I typically will take any new drive and do some type of stress test [1] on it for 18 to 24 hours to see if it fails with that initial constant use.
[1] Constant reformatting for example writing 0's to the entire disk 7 times etc.
Do you have the figures for raw losses including at burn-in?
I mean, sure, you burn-in and do warranty returns so buying Hitachi would still seem better - but if one needs a drive to just work then it's key to know the overall failure rate.
Why not take apart some failed drives and see what you can find out as to why they failed vs. ones that did not. Perhaps info that might be of interest to the manufacturer but in any case would make an good blog post. Or maybe you can cannibalize and use again.
[1] I did an "air crash investigation" on a RC Chopper that had crashed. I found out that a servo failed because a part in it was plastic (a gear) (as opposed to, I think, brass). Consequently the jerky move that I made was enough to cause the plastic gear to loose a tooth and then I lost control.
"<a href='https://www.backblaze.com/blog/hard-drive-reliability-update... src='https://www.backblaze.com/blog/wp-content/uploads/2014/09/bl... alt='Hard Drive Failure Rates by Model' width='560px' border='0' /></a>"
should be "width='560'" not "width='560px'"
Most ofther file-servers have a front-facing drive caddy, that usually has LEDs on the front to indicate disk access or errors. This is great because you can walk into the datacenter, and SEE which disk has failed. With the backBlaze system you can get /dev/DriveID but not know where in the array that particular disk is.
Second to that, you probably don't want to go single-source for your drives -- maybe use Hitachi with a mix of WD.
Can you explain why not?
Could also be caused by bad grease, or shoddy bearings, or pretty much anything.
There is security in diversity.