Hard Drive Reliability Review for 2015
backblaze.com
backblaze.com
"A relevant observation from our Operations team on the Seagate drives is that they generally signal their impending failure via their SMART stats. Since we monitor several SMART stats, we are often warned of trouble before a pending failure and can take appropriate action. Drive failures from the other manufacturers appear to be less predictable via SMART stats."
~10 years ago, I remember google research put out a highly cited paper wherein they found that SMART stats were not a particularly strong indicator of impending drive failure (50% of drives had no SMART indications of problem before failure). http://research.google.com/pubs/pub32774.html
Has this now changed (at least for Seagate)?
Reliability/longevity is nice but a signal of impending failure is far more valuable from an operations point of view.
Many SMART stats aren't particularly useful in predicting failure as they simply correlate to the age of the drive in some fashion.
Also, here is our data on every single SMART stat for all of the drives we have: https://www.backblaze.com/blog-smart-stats-2014-8.html
Gleb (CEO, Backblaze)
First thanks for all your company's sharing of such data, as well as the pod open platform. Kudos.
Second, can you do a Writeup specifically and only about SSDs?
Thanks
It's probably not 5% of our total storage, though.
* even the slowest SSDs have significantly higher I/O rates than the best mechanical drives, and the comparison between best-in-class mechanical and enterprise-class PCIe SSDs is just ridiculous: a 15K SAS drive will do 200 IOPS, a high end SSD will do a million
* 15K SAS drives will top out around 250MB/s on bulk sequential reads (that's a best-case scenario), high-end PCIe SSD are in the 2.5GB/s range
* HDDs have a latency of 10~20ms, SSDs have a latency of 100~200µs (RAM has a latency of ~100ns)
Have you productized these learnings in a a powertop-like tool for Linux?
Smartmontools are not intuitive enough for the layman to use in any meaningful way.. And backblaze has really built some serious learning here that could be of use to everyone.
(I say this as I've been recently built a freeNAS box with a combination of Seagate NAS and WD Red HDDs - the WD's make it easy to look at the smart stats and know what's going on. The Seagate ones, not so much.)
I suspect rotating drives have a variety of several failure modes, some of which could be predicted by SMART, others which it's unlikely to be predicted.
Each new model is probably bound to have a different pareto of failure modes.
Also, the fact that backblaze are publishing most of their data online is very cool.
Very unfortunately, HGST has apparently scaled back Deskstar sales and development significantly since the acquisition. I guess it has to do with WD selling off some of HGST's 3.5" assets to Toshiba in order to appease competition authorities. See also https://news.ycombinator.com/item?id=10057519
"In May 2012, WD divested to Toshiba assets that enabled Toshiba to manufacture and sell 3.5-inch hard drives for the desktop and consumer electronics markets to address the requirements of regulatory agencies."
[1] http://seekingalpha.com/article/3827746-seagate-western-digi...
Allegedly Cirrus Logic controller had a manufacturing defect and died due to heat. Myself I always suspected that very peculiar and strong smelling rosin flux. PCB was drenched in it, this type of flux is usually highly activated and requires cleaning, otherwise acid will eat solder joints and copper away, especially in humid and hot environments.
http://www.theregister.co.uk/2015/10/19/mofcom_says_yes_wd_h...
Of course, the reality is China wanted a piece of WD, and used the merger as leverage to get it. I would expect by 2017, HGST drives are just as shit as WD. Which is unfortunate, because the Japanese designed one hell of a hard drive.
There is probably some good news in the article though, for what it's worth: "At that time John Cyne ran WD. He since retired, with HGST bss Steve Milligan taking on his job." ... "In other news Western Digital has announced a new executive management team, and it looks almost like HGST executed a reverse take-over of Western Digital." ... "A person who was close to the corporate action in Western Digital and HGST said: 'All key positions are with HGST people; it's a reverse buyout. First HGST took Coyne's money to buy themselves (probably with a clause that Milligan is becoming CEO) and then they watched WD dismantling itself.'"
I had one of the IBM Deskstar (aka. Deathstar for the high failure rate). IBM sold the HDD business to Hitachi, who sold it to WD.
And now, Deskstars are as reliable as they come.
I believe these Deskstar types and derivatives are the same essential mechanism and processes as those old quantum drives (probably especially in terms of the QA processes). The heads and the technology have improved to give better capacity of course, but I've been buying and relying on these drives for ~25 years at this point.
I'm not surprised to see them showing up well on these charts.
I have been relying on these drives for 10+ years too and hope Toshiba keeps producing them.
[1] http://seekingalpha.com/article/3827746-seagate-western-digi...
> Quantum [...] were beating IBM and IBM needed more capacity so IBM bought them.
No, Maxtor bought Quantum's HDD business, and Seagate later bought Maxtor [1], [2], [3].
> Western Digital who sold the drive business to Toshiba
No, WDC sold some assets related to desktop (not server) 3.5" (not 2.5") HDDs to Toshiba [4], [5], [6].
[1] https://en.wikipedia.org/wiki/Maxtor#Acquisition_of_the_Quan...
[2] http://www.cnet.com/news/maxtor-buys-rival-quantum-to-become...
[3] http://www.pcmag.com/article2/0,2817,1966943,00.asp
[4] https://en.wikipedia.org/wiki/HGST#History
[5] http://techreport.com/news/22553/toshiba-becomes-third-playe...
([#drives][# failures]) / [operating time across all drives]
Wat? The numerator and denominator seem unrelated. What is being measured here?
To me, it would make more sense to look at time to failure. Together with data on the age of the drive and the proportion of failures each year one could create an empirical distribution to characterise the likelihood of failure in each year of service. That would give a concrete basis from which to compare failure rates across different models.
With right censored data (as this is), if you measure age at death but then you're only modelling already failed drives, so you'll under represent good drives.
It would be good to see some statistics done so we can see confidence intervals around a hazard rate at different ages.
It is raw data that they provide in an excellent fashion.
You can consume it raw and trust their high driver numbers to drive p high enough for you, or you can use their data for real science. either way, it is a great deal that they take the time to share it.
It's all just a scaling: you have a number of broken drives in a corner of the datacenter in the wire bucket that says "broke during 2015", you count them, divide by total hours of that type of disk running (since they may have been brought in commission at different points), and then scale it so you get it in percent-per-year, not likelihood-per-hour.
It smells of someone explaining code, rather than illustrating an important engineering formula, but there's nothing wrong with the rescaling calculation per se.
Perhaps the problem is the specific example given. 100 is the size of the drive fleet and also the multiplier required to convert to percentages. Let's assume you are right and the 100 in the equation is not #drives.
Even so, I find the approach questionable. If the point is calculate the proportion of failures then that (overly simplistic) calculation is:
[#failures] / [ #drives] = 5 / 100 = 5% failure rate.
But this isn't what's calculated. Instead the author calculates the proportion of drive-years per annum affected by failure. For the 100 drives in the example the cumulative number of operational hours given in 2015 is 750K hours (out of a possible 876K hours, had the drives been operating 100% of the time).
That's a problem because 750 / 876 = 85.6% of total time.
5 / 85.6 = 5.84% "failure rate" which seems to me an overstatement.
The problem gets worse as the number of operational hours decrease. Imagine for a moment the 100 drives only operated 50% of the time in 2015. We have:
100 (5 / ((875K*0.5) / 875K)) = 10% "failure rate". This despite only 5% of the drives having failed.
Wat?
This would determine if the failure rate was constant for the life of the drive (meaning random failure) or is it age related (infant mortality or old age).
25 drives that fail after 1 week plus 25 that fail after 50 weeks is different to 50 drives that fail one per week.
As one does.
Like this, for example: http://www.tweaktown.com/articles/6028/dispelling-backblaze-...
I remember a study Google did on harddrive reliability and it seemed to show that heat had little to no effect on it. I also don't regard consumer-level as being a bad thing. As a consumer, I kind of want to know which drives are built for abuse better. All drives fail; which drives fail more and at what cost?
It's also a bit disingenuous to criticize back blaze's methodology when you know that a 'comprehensive study' under more controlled conditions will NEVER actually happen with the necessary sample size to draw conclusions.
Stress testing is a valid methodology for determining reliability - eg car makers crash their cars into walls at high speed to make sure they are safe, or use a robot to push the brake pedal a million times to see when it fails - so they hardly deserve criticism for pushing the drives hard. More information for the consumer is a good thing.
1. They rip hard drives from external enclosures.
2. They have too much vibration in their pods.
3. They don't correct for temperature.
4. They worked the drives too hard.
The whole article reads like the excuses of someone with a vested interest in discrediting evidence of their favorite brand's poor performance. I don't think the take away from the data provided by Backblaze is "I can expect to get a failure rate of exactly 1.231971 if I buy brand X's hard drives." The end-user-useful conclusions are things like "HGST's drives are the best," and "6TB drives are less reliable than 4TB drives right now."
Sure, all of the factors listed in the criticism may play a role in the failure rates (except the external enclosure bit, since A. The majority of the "shucked" drives were 3TB, and B. They've outgrown that practice.) But they only have the weakest of justifications for believing that those factors vary systematically across the manufacturers. And indeed, even those factors did vary systematically we'd still get the right answer if we had made the more general conclusions. For example, if the vibrations in the seagate-only enclosures are greater than the vibrations in the HGST-only enclosures, that can only be because the HGST drives are better and vibrate less. Or alternatively, maybe the pods all vibrate the same, but HGST is better because it is more resistant to vibrations.
I regularly buy external HDDs, rip them out and put them into desktops and laptops, put them back in different enclosures, and so on. As a result, my HDDs experience a lot of movement and extreme temperatures (e.g. being left in the trunk of a car on a hot summer day). It's good to know which models are the most likely to survive such abuse in the long term.
https://www.backblaze.com/blog/farming-hard-drives-2-years-a...
And the discussion on HN:
That story is linked down near the end of the article:
On the other hand, those who do remember IBM hard drives probably remember them as the Deathstar, so HGST might not want to be associated with their old home so much.
But actually I don't see how this makes them biased in any way. All drives essentially sell for the same amount (and Amazon pays a percentage of that) so if you trust the info as being accurate (and why wouldn't it be?) then how could it biased then given there is such little lattitude in pricing?
And who is going to accuse them anyway? People who read HN? If so, so what?
The data presented is a nice shortcut answering the question of "which drive should I buy" without having to read all of the charts and most importantly think.
Lastly, you don't have to buy from amazon just because they give you a link but it does make it easier to see a price and compare to whatever vendor you might typically use (or provide several links to different vendors).
As a backup company, we hold ALL our customers data, so our reputation is incredibly important to us. People MUST trust us as impartial and trustworthy and not sleazy or we would go out of business quickly.
1) So what does it look like now with what you are doing? For example you are offering free credible information about drive reliability which contradicts what you actually do which is make using drives for backup irrelevant. While I am sure that the following is not the case, I could easily say that you are doing this to make people think drives aren't reliable and hence they need backblaze! Wow look at drive failure I should DIY this! (Do I think that is your strategy? To repeat I don't..)
2) Note that http://www.dpreview.com was purchased by Amazon and it has only grown larger and more reputable (in terms of the reviews) since then. And they openly link to Amazon and they could easily be accused of a tremendous bias but apparently they either aren't worried about that or the effect is nominal.
3) I can fully understand, as a business decision, why you might not want to "cheesy" up (my words) your site with amazon links or perhaps you might feel the 3% is not consequential enough to do so. It is certainly a judgement call. However don't assume that everyone that would be a potential user of your company really would think that way because I can assure you that isn't the case.
> we hold ALL our customers data, so our reputation is incredibly important to us.
The fact that you are earning money from affiliate links does not mean you are not reputable and doesn't give me any less confidence that my data will be safe. It's a non issue (for that reason). You have a right to earn money in any reasonable fashion. Affiliate links are an accepted way to earn money (we aren't talking about selling customer data). If anything I think almost the opposite. I want to know that you are making money and robust in business practices so you have the funds to insure your operation will continue for the foreseeable future.
We struggle with it internally, I assure you we doubt ourselves all the time. :-) Some companies have an "informal fun loving" outward appearance, like if you purchase from Zappos they send emails like "the magic elves are making your shoes, we will send them along very soon..." But bankers tend to wear suits and ties and appear "very serious" in their communications even while frittering away your money on sub prime mortgages.
Anyway, the point is I'll forward your note along and heck, maybe next quarter our drive stats blog post will have Amazon links and we'll make a little extra money. :-)
However it's important that you wrap this in the proper words [1] not just plop the links on the page.
You need to explain the links but without apologizing for putting them there. You can even say perhaps that you were asked to do this (because you were). And don't chicken out and say you are donating the $$ to charity or anything like that.
Depending on how you write this, you will minimize the whiny blowback (if any). That said, running a business is not running a popularity contest to the tune of the most vocal commenters on HN or reddit or wherever.
If you are not doing so already you might want to issue traditional press releases with your results as well.
Of course if you do the links (and I would try this for more than one quarter) if it works or if it doesn't work you can then do a blog post on that!
[1] In the business I am in we charge for a service that our other competitors give away for free. By wrapping it in the proper words we often get a thank you instead of a complaint.
Considering a 4 TB hard drive has to track 32,000,000,000 individual bits, allowing reading and writing repeatedly of each one, on platters that are spinning 120x per second, spaced a hair's width from their heads...I think it's actually incredible.
As for SSDs, we keep wishing that we could switch to them, but they're still 10x more expensive on a $/TB basis. That may change in the next few years, and if it does, we'll look forward to sharing data on SSD usage at scale as well.
Gleb (CEO, Backblaze)
Thanks again for sharing the drive reliability statistics.
I don't mean to insult, just to ponder the relevance of such long-term studies on tech that changes so quickly.
Large companies may buy Seagate due to the price advantage and the fact that their storage systems can better handle the drive failure rate.
6TB 1.89% 4TB 2.19/2.99% (depending on model) 3TB 5.1/28.34% (depending on model) 2TB 10.1% 1.5TB 10.16%/23.86% (depending on model)
For example, for less reliable manufacturers there might be a "if you get past first N weeks, you are fine" pattern, or a failure cliff exaclty 1 week past the warranty period, or something equally entertaining.
I've got 5 Western Digital drives which have failed out of original purchase of 6. Now I'm wondering if it's really worth it trying to go through the RMA process (I need to figure out exactly how old they are and how long the warranty is) or if I should just give up on Western Digital and go with a different manufacturer... though I am not looking forward to spending that amount of money all at once.
I've had it for four years now and there are no warnings of any kind yet, so I guess I got one from a good batch.
Amazing the 4TB hitachi with twice the platters of the 2TB fail less.
(and I will never buy seagate again for home pc or servers, even before this report I could have told you they are unreliable)
The Seagate 3TB were awful, but their 4TB seem to be just fine.
I cracked open the enclosures and the drives are just fine. I still use them for backups with no errors.
https://www.wish.com/search/2%20tb%20thumb#cid=5683434cce922...
Beyond that, flash drives tend to have low write durability and horrible performance on large writes (because of poorly implemented garbage collection).