Backblaze Drive Stats for Q1 2024
backblaze.com
backblaze.com
For example, in the Pro-sumer space, both WD's Red Pro and Gold HDDs report[1] their endurance limit as 550TB/year total bytes "transferred* to or from the drive hard drive", regardless of drive size.
[1] See Specifications, and especially their footnote 1 at the bottom of the page: https://www.westerndigital.com/products/internal-drives/wd-r...
I ended up buying a similar sounding but not same model from CDW.
The Amazon refurb drives (in this class) typically come with 40k-43k hours of data center use. Generally they're well used for 4½-5yrs. Price is ~30% of new.
I think refurb DC drives have their place (replaceable data). I've bought them - but I followed other buyers' steps to maximize my odds.
I chose my model (of HGST) carefully, put it thru an intensive 24h test and check smart stats afterward.
As far as the 5yr warranty goes, it's from the seller and they don't all stick around for 5 years. But they are around for a while -> heavy test that drive after purchase.
If you buy used, you're avoiding the first form of failure.
GoHardDrive is notorious for selling "new" drives with years of power on time. Neither Newegg nor Amazon seem to do anything about those sellers
https://news.ycombinator.com/item?id=39262314
>https://listofdisks.com/ https://www.cpuscout.com/ https://diskprices.com/ https://gpuprices.us/ https://instances.vantage.sh/ https://tvpricesindex.com/
Like that bookstore that just happens to retail some stuff too.
When they branch out to selling everything including fresh vegetables, motor oil, and computing services, then maybe they might be more comparable to the overgrown bookstore.
They sell other stuff too but they’re still pretty photo and video-centric, laptops notwithstanding.
That's a milestone. Imagine the racks that were eliminated
I'm imagining about 3/4ths ;)
But they run many drive types.
Right after launching B2, in late 2015, they made their post about storage pod 5.0, saying it "enabled" B2 at the $5/TB price, at 44 cents per gigabyte and a raw 45TB per rack unit.
In late 2022 they posted about supermicro servers costing 20 cents per gigabyte and fitting a raw 240TB per rack unit.
So as they migrate or get new data, that's 1/5 as many servers to manage, costing about half as much per TB.
It's hard to figure out how the profit margin wasn't much better, despite the various prices increases they surely had to deal with.
The free egress based on data stored was nice, but the change still stings.
Maybe I'm overlooking something but I'm not sure what it would be.
In contrast the price increases they've had for their unlimited backup product have always felt fine to me. Personal data keeps growing, and hard drive prices haven't been dropping fast. Easy enough. But B2 has always been per byte.
And don't think I'm being unfair and only blaming them because they release a lot of information. I saw hard drives go from 4TB to 16TB myself, and I would have done a similar analysis even if they were secretive.
Inflation is not even close to that level.
And those hardware costs already take into account inflation up through the end of 2022.
That would work if they fully recouped the costs of obtaining and running the drives, including racks, PSUs, cases, drive and PSU replacements, control boards, datacenter/whatever costs, electricity, HVAC etc. and generated a solid profit not only to buy all the new hardware but a new yacht for the owners too.
But usually that is not how it works, because the nobody sane buys the hardware with the cash. And even if they have a new fancy 240TB/rack units, that doesn't mean they just migrated outright and threw the old ones ASAP.
So while there is a 5x lower costs per U for the new rack unit, it doesn't translate to 5x lower cost of storage for the sell.
You can look at their stats and see that the very vast majority of their data is on 12-16TB drives, and most of the rest is on 8TB drives. Even with those not being the very newest and cheapest models, their average server today is a lot denser and cheaper than their brand new servers 8 years ago.
I on the other hand have a 4U 48 bay Chenbro so drive failures are somewhat significant for me lol.
Redundancy wise it's 4 raidz2 vdevs with 12 drives each and backed up to rsync.net I have had 2 drive failures, one was shortly after commissioning and the other happened a few months ago which was pretty random.
I'm using HGST drives, specifically 8TB He8 and they have been really solid in operation since 2016. I don't have any spares left now though so when I get back to where the chassis is hosted I will be doing a rebuild onto 16TB drives.
On the other hand in my professional life I experienced arrays that had multiple drives fail in quick succession (especially around 2010-2012 era) from less ... reliable brands cough Seagate cough.
So I would consider 2 failures from ~1.5M drive hours to be very good and thank Backblaze for convincing me to shell out on these rather more expensive drives.
1. 2 failures across 1.5M hours is something around a 1.2% AFR, which is good, but not significantly below Backblaze's average. Definitely better than the stats on their worst drives.
2. Assuming some premium for higher-reliability drives, and some required storage growth over time, the most efficient drives to buy are those that fail at exactly the rate that lets you replace them at with higher-capacity drives as needed for storage growth. I'm personally at the point where I'm decommissioning 2TB drives with 100k hours; I'd be better off having saved some money and having the drives fail now.
https://www.backblaze.com/blog/ssd-edition-2023-mid-year-dri...
Well, unless you're putting large numbers of consumer SATA drives into massive storage arrays with proper power and cooling in a data center.
In my experience, the drives report "healthy" until they fail, then they report "failed"
I've personally never tracked the detailed metrics to see if anything is predictive of impending failure, but I've never seen the overall status be anything but "healthy" unless the drive had already failed.
> I've personally never tracked the detailed metrics to see if anything is predictive of impending failure
Backblaze has!
From experience, we have found the following five SMART metrics indicate impending disk drive failure:
SMART 5: Reallocated_Sector_Count.
SMART 187: Reported_Uncorrectable_Errors.
SMART 188: Command_Timeout.
SMART 197: Current_Pending_Sector_Count.
SMART 198: Offline_Uncorrectable.
That's good to know, I might start tracking that. I manage several clusters of servers and hard drive failures just seem pretty random.However, one time a drive got a burst of reallocated sectors, it stabilized, then didn't have any problems for a long time. Eventually it wouldn't power on years later.
I've never seen an HDD fail overnight without any indication at all.
Are most drives retired without failing?
I'd expect so given that HDDs are still having significant density advancements. After a while old drives aren't worth the power and sled/rack space that could be used for a higher capacity drive. And, yeah, it makes these statistics make more sense together.
Edit: plus they are just increasing drive count so most drives haven't hit the time when they would fail or be retired...
Yes, certainly.
One can watch both SMART indicators as well as certain ZFS stats and catch a problem drive before it actually fails.
I like to remove drives from zpools early because there is a common intermediate state they can fall into where they have not failed out but dramatically impact ZFS performance as they timeout/retry certain operations thousands and thousands of times.
I'd like to see instead something like mean time until 2% of the drives fail. That'd actually be comparable between drives. And yes, it would also mean that some drive types haven't reached 2% failure yet, so they'd be shown as ">X months".
This is what a Kaplan-Meier survival curve was meant for [0]. Please use it.
Also, it'd be great to see the confidence intervals on the annualised failure rates.
[0] https://en.wikipedia.org/wiki/Kaplan%E2%80%93Meier_estimator
(in reality they'd probably have failure rates spike at some point, but the idea stands. And they explicitly said they retired a bunch of 4TBs)
So you have usually a lifetime of drive tput and start/stop values you want to stay under, and depending on how accurate your data is for each drive you may push beyond the drive warranties. But you will generally stop before the drive actually fails.
While they seem to get retired, it's not as quick as we'd think.
They go on to list 3 Seagate models that share one common factor: Sharply lower drive counts. Backblaze had a lot fewer of these drives.
All of their <5 failures are from low quantity drives.
I have confidence in the rest of their report - but not with the inference that those 3 Seagate models are more reliable.
This video presents AFR, failure rates, derived from prior backblaze reports, aggregated.
Definitely worth a watch if you're interested in this report.
Seagate continues to trail behind competitors.
I guess they're basically competing on price? Because with data like this, I don't know why anyone running data center would buy Seagate over WD?
Probability is a strange thing, yo. The odds of a specific person winning the lottery are effectively 0, but someone's going to. Looks like I've won the "WD means Waiting Death" sweepstakes.
SMR takes advantage of the fact read heads are often smaller than write heads, so it "shingles" the tracks to get better density. However, if you need to rewrite in between tracks that are full, you need to shuffle the data around so it can re-shingle the tracks. This means as your array gets full or even just fragmented, your drives can start to need to shuffle data all over the place to rewrite a random sector. This does hell to drives in an array, which a lot of controllers have no knowledge of this shingling behavior.
Shingled drives are OK when you're just constantly writing a stream of data and not going to do a lot of rewriting of data in betweeen. Think security cameras and database backups and what not. They're complete hell if you're doing lots of random files that get a lot of modifications.
https://www.servethehome.com/wd-red-smr-vs-cmr-tested-avoid-...
Either way it made me never want to use WD for drives in arrays and not trust their labeling anymore. "WD Red" drives lost all meaning to me; who knows what they're doing inside.
I'm not ruling that out. The whole debacle was so amazingly tonedeaf that I wouldn't be surprised if they did that behind the scenes. I wrote this at the time: https://honeypot.net/2020/04/15/staying-away-from.html
If a brand sells bad drives, they should be aware of the reputational damage it causes. Otherwise there is no downside to selling bad drives.