Backblaze Drive Stats for 2022
backblaze.com
backblaze.com
This means that these stats may not be a very meaningful indication when purchasing a few drives from a retailer.
Never the less, the world is full of situations where you have to make choices with insufficient data. I'm still going to prefer to avoid Seagate hard drives where practical.
Seagate is the manufacturer with drives that fail at 10x the rate of the competition. I'll avoid, because drive failures are annoying.
The other way you could read the comment would be that the stats don't mean much since Backblaze can tolerate higher defect rates if they negotiate a low enough price. But to me, that doesn't make sense since it wouldn't affect the failure rate statistics at all anyway.
As a consumer, you're unlikely to keep an in-box spare, so even if you have a good backup system, a drive failure is still going to be an inconvenience, and leave you without redundancy until your replacement is shipped to you, so you might pay 20% more for drive A.
It is because of theses statistics that you have (at least an inkling of) a fighting chance to pick the drive that more suits your usecase. Or perhaps a notion of brand performance.
Of course, you won't get statistics about drives that they don't buy. And if you then assume that more expensive drives have better failure rates then that might skew the data. But in there lies a lot of assumptions you really can't make. For one, that an expensive drive for a consumer is also an expensive for backblaze. And also, that there is correlation between end user price and failure rates.
Which generally I don't really think there is, within the alternatives that the consumer have access to (and reasonable price-ranges) at least. And it probably will vary between different markets as well.
The top level commenter points out that this strategy may not make as much sense for the individual consumer where you don't have backblaze levels of drives to amortise the failure probability and deal with it.
That's it, that's the point. It's not a rebuttal of the backblaze post at all, nor is it intended to be. The top level commenter was simply pointing out "don't buy mostly seagate drives because backblaze does" as backblaze explained their strategy in drive selection and some people may not consider the differences between their use case and backblaze's.
Top level comment: >This means that these stats may not be a very meaningful indication when purchasing a few drives from a retailer.
That doesn't follow.
You assume people read the fine print and consider fine details.
The vast majority of Joes are going to ask who/what Backblaze is and what they're buying most and buy that, screw the details.
Given that the manufacturer knows that Backblaze publishes its influential drive stats... do we know that the drive units that the manufacturer ships to Backblaze are representative of the quality of the same models available at retail?
> Do we know if these drives are the same as you would purchase in a retail outlet
Seagate (for instance) won't sell ANYBODY drives directly, so they force us to get bids from various resellers and distributors. So when we pick the lowest price, I don't think Seagate knows who the drives are going to, but there might be a trick in there somewhere I am not aware of.
Every drive must go through a series of tests before it is judged suitable for sale to the consumer. Drives that fall below a certain threshold are judged 'failed' but there can be a variety of drives that fall within the 'passed' range. Some are just barely above the threshold, others are way above.
It is kind of like binning for CPUs. The very best ones can have their base clocks increased and sold at a premium. The same kind of differentiation could be done for hard drives.
They also definitely did a binning process where all drives being sold to certain manufacturers would be heavily tested for certain types of failure modes and only those that passed the most rigorous tests in a certain area would be approved to be sold to certain manufacturers.
Of course, each manufacturer had their own things that they were concerned about.
I don't know if they still do binning today, but they certainly did it back in 1989.
I ran an array with 1.5T and 3T Seagates and within a 1-3 years, I replaced at least 1/3 of my Seagates, but I made sure to replace the Seagates with Hitachis and Western Digitals, even though they were slightly more expensive.
I don't see so much significant higher failure rate on my 12TB Seagate drive array.
That being said with 50TB drives just around the corner [1] and SSD caching becoming the norm, 12 bays consumer NAS will become rare.
[1] https://www.anandtech.com/show/18733/seagate-confirms-30tb-h...
My iPhone shoots 90 megabyte photos.
Specifically, the night mode on recent (2y) devices.
If it takes days to fill your new 50 TB backup drive, you have the same problem RAID5 has with rebuild times. The drive might fail in the middle.
RAID isn't a backup mostly because if you overwrite RAID data, you wrote over every copy at once. Non RAID offsite backups don't solve any problem related to drive size, they just make it a lot less likely for a single event to take everything out in the same minute.
50TB at 200MB/s is about 72h. Doesn't seem to be a particularly problematic rebuild time (and that's assuming 100% filled). Of course you need to do regular data scrubbing. If your rebuild is the first time the data is being read in 5y, that might not go so well.
Assuming the device is offline for users during that time, otherwise there may be reduced throughput.
If you have standby sets (like a read replica of a db that can take over for master, or a DR site), you can switch temporarily.
It’s good to know what throughput you have under load for this reason among others.
Also, you can slow or pause some RAID during peak hours.
RAID 1 joins the chat
Er what? I've had production machines lose a drive, have it replaced, and never leave production. Why would you need to restore from backups?
You could argue to remove the driver that are represented less than 100 times (and it will be a stretch).
But once you have so many samples, there is no reason to not believe that your driver will not follow the same distribution.
Out of the many hard drives that I've owned in my life, the only ones to die on me have both been Seagates, with the first one being the infamous ST3000DM001, a hard drive so shit it has its own Wikipedia article (https://en.wikipedia.org/wiki/ST3000DM001).
AFAIK there hasn't been much of a difference shown between consumer workloads and enterprise workloads for HDD lifespan either, it's just cope and theorycrafting from people who are emotionally attached to the idea of Seagate not being shit for some reason.
No other drives seem to have such problems with being used in this fashion: what is your theory for why Seagate drives are uniquely affected by being in the pods in some fashion that would not also affect WD drives or HGST drives or whoever else? Are WD drives not affected by vibration for some reason? For a while Seagate Deniers latched onto the first-gen pods as maybe being the answer but they're all long gone at this point, this failure-rate anomaly is continuing even in the newer pods.
The only reasonable possibility would be that Seagate drives are simply constructed in an entirely different, less resilient fashion, which (a) is not factually supported in any way afaik, and (b) would still be very relevant for consumers to know! It's not like a home PC is vibration-free either after all.
There comes a point when it's not "steelmanning" it's just denial of reality in the face of consistent evidence. Like you're not "steelmanning" climate change you're just a denier.
The data has pretty consistently showed the same thing for 10+ years. It's not a "random sampling bias" that uniformly affects everyone except Seagate in the exact same way almost every single survey, it's not some magical factor that makes Seagate drives uniquely unsuited to storage-array usage but magically resilient when used in a home PC, it's not first-gen backblaze pods being bad, it's just Seagate putting out shitty drives, period the end. All these extremely complex theories to get around the very simple conclusion that Seagate has shitty parts or shitty QC and the failure rates are slightly higher as a result.
A lot of the Seagate models are relatively OK, but almost all of the "outlier" drives with really high failure rates are Seagate. It is the old bayesian probability thing: get rid of Seagate and you've gotten rid of almost all of the models with high failure rates.
edit: sorry Samsung on the brain since they have another wave of SSD failures too /laugh
The first to fail was a Miniscribe 5¼" 20MB, it died after a week. I replaced it with a 3½" Seagate connected to an Adaptec RLL controller - this being before IDE was a thing the controller decided the encoding scheme and RLL gave you 50% extra storage space. This Seagate never died on me, I probably have it around somewhere still. Then came the Western Digital IDE drives, two of those died and took ~250MB of data with them to data heaven. They were followed by another WD, this one 1.2GB - it died. When I built a system around the famous Abit BP-6 motherboard I put in two Maxtor drives - which I should not have done, one of them died within 2 years. Meanwhile I'd helped my father get a new drive, one of those fancy IBM Deskstars. Within a year it had turned itself into a Deathstar like most of its brethren seem to have done. Then there was the 10 GB Travelstar which died, then the 20GB Travelstar which also died - but survived the canoe expedition over the Yukon safely tucked away into the water- and probably bulletproof solar-powered Virgin Webplayer I had made to document the 3½ month trip - and the 20GB Toshiba. The 2TB WD Green, dead. One of the 1TB WD Greens, dead - but its mate still running strong at a 120.000+ running hours with its 2TB brand mates coming in second at 100.000+ hours. In the DS4243 array I have replaced 5 15K 600GB drives, good that these were cheap as dirt (but take quite a bit of power to run, hence the low price).
Notice that I have not had a Seagate drive die on me yet. There are a few in an array somewhere here but those still do their job even though they're quite old by now. Maybe I'm just lucky in that I never bought any of the failing types since these problems seem to be related to specific types.
I'm wary of Amazon due to co-mingling issues. I can't be sure I'm not getting something incorrectly packaged/labelled by another seller so I get a reconditioned (or simply counterfeit) unit instead of a new one.
Update: This issue has been covered in the Blackblaze 2020 report. They apparently existed in parallel. The original HGST drives, while still existed as of 2020, were being gradually phased out...
> These drives obviously share their lineage with the HGST drives, but they report their manufacturer as WDC versus HGST. The model numbers are similar with the first three characters changing from HUH to WUH and the last three characters changing from 604, for example, to 6L4. We don’t know the significance of that change, perhaps it is the factory location, a firmware version, or some other designation. If you know, let everyone know in the comments. As with all of the major drive manufacturers, the model number carries patterned information relating to each drive model and is not randomly generated, so the 6L4 string would appear to mean something useful.
I believe hese new HGST drives are WDs.
The HGST drives in the older backblaze stats (which showed good reliability) continue to be manufactured by/evolved into Toshibas. There are several HGST model numbers that remained in Toshiba's lineup.
https://www.anandtech.com/show/5635/western-digital-to-sell-...
Yes. WD did a rebranding exercise to try and get rid of HGST and UltraStar, not only did the retail channel and wholesale channel backfire, a certain large enterprise customer insist on the same HGST and UltraStar HDD, same model number, basically same everything. While HD could certainly ignore retail and wholesale, they cant ignore enterprise. So you still see UltraStar, and in many cases Wholesale are still calling them HGST.
Helium atom is too small and leaks through everything, eventually.
Really hope I am proven wrong.
Unfortunately what backblaze is excellently documenting is not archival use.
I've got half TB and 1TB WDC drives that are over a decade old and still spin up fine, single platter and run cool and quiet even air filled.
I think 4TB is the cutoff for air-filled but not sure anymore.
Suppose the drive were encased in solid hydrogen? Hydrogen freezes at 14 K and helium boils at 4 K so there's a 10 K range where you could have both solid hydrogen and gaseous helium.
Hydrogen atoms are bigger than helium atoms, but what matters is the gaps between the hydrogen in the solid. I was unable with a bit of Googling to find out anything about the physical structure of solid hydrogen.
If it becomes a vacuum or replaced by air, helium drives will fail.
Installing their linux app was out of bounds of normal procedures, but it be what it b2 <hangsHeadInShame>
Anyway, I haven't seen any apparent data loss over the last few years. I'm considering doing a full restore, just for the heck of it. (Especially since this article convinced me I should start replacing the aging drives in my NAS!)
At no point has it risen above 1MB/s when backing up to S3 (Backblaze) when other methods routinely saturate my upload (40Mb/s)
(I've assumed TiB over TB given they are the sizes reported by the drive manufacturer.)