Backblaze Hard Drive Stats for 2016
backblaze.com
backblaze.com
What you're seeing is not the most reliable, but the most reliable under their particular set of bad conditions. The reliability under your circumstances might be wildly different.
Their pod approach even results in worse performance: more vibration results in more effort spent by the drives trying to keep the head positioned in the right places, increasing latencies and decreasing read performance. The higher density they push for with their pods actually compounds this. Drives vibrate. Fans vibrate (chassis, PSU, CPU, the works) . If you don't pad them all appropriately, the vibration of the components ends up causing the entire rack to vibrate. By the time you've got several dozen of drives in the rack, the whole rack will be feeling the effect.
[1] http://www.dell.com/us/business/p/poweredge-r530/pd?ref=PD_O...
Storage servers do tend to do stuff to dampen vibration, but then it's a question of storage density. When you've got something like a backblaze pod that's 60 drives in 4U of space. It'd be more unusual to have 60 drives in total in a rack full of normal servers. Now consider 60 drives / 4U, and you'll have multiple of these storage pods in a rack. It adds up.
You do have a valid point but I also wanted to add that my harddrive failures in my gentle home environment matches BackBlaze's more hostile environment. I have 50+ harddrives in various capacities from 2TB to 10TB and my differences from BB include drives being iso-mounted on rubber grommets and they also do not run 24/7. (I only power them up when I need archived files off of them.)
Even with those stark operational differences, my harddrive failures exactly match the generalization of Backblaze's data: More failures with WD Green and Seagates especially the 3TB drives. My HGST drives purchased over the last 4 years have zero failures so far.
With consumer drive outside the case, on a bench: http://broadley.org/bill/consumer-no-vibration.png
Inside: http://broadley.org/bill/consumer.png
Server drive inside: http://broadley.org/bill/server.png
We went with the RAID edition.
With respect to vibration: we found vibration caused by adjacent drives in some of our earlier drive chassis could cause off-track writes. This will cause future reads to the data to return uncorrectable read errors. Based on Backblaze's methodology they will likely call out these drives as failed based on SMART or RAID/ReedSolomon sync errors.
Yev from Backblaze -> In Storage Pod 6 we actually did a lot more to work on the dampening (and 5.0 as well, but 6.0 builds on it). The drives are all placed in to guided rows and are held down by individuals lids w/ tension on them. You can take a gander at the latest build here -> https://www.backblaze.com/blog/open-source-data-storage-serv...
No. I've seen this argument raised a few times whenever hard drive failure rates comparing manufacturers has occurred. (Previously on HN: https://news.ycombinator.com/item?id=7119323 ) Backblaze puts more stress on their drives than typical users, but a drive which works well in their environment is also very likely to do the same in a less stressful one. This is how accelerated life testing is done.
https://en.wikipedia.org/wiki/Accelerated_life_testing
Their pod approach even results in worse performance
Not once in the data presented in the article is performance mentioned. This is not about performance, it's about reliability. It may be anecdotal, but the numbers I see correlate well with my experience and many others' experiences I've heard of.
More amusingly, I remember coming across a forum discussing data recovery, which happened to be split by manufacturer, and noticing the Seagate forum had an overwhelmingly large number of topics containing people asking for help with their dead disks relative to the WD, Toshiba, Hitachi, etc., which was disproportionate to their marketshare.
Thanks for that, I guess.
A few years ago, the consensus (and data) was that they were mostly the same thing and it was arguable to pay double for the enterprise one.
But... in recent years, the consumer market have moved toward ever cheaper, slower and power savings hard drives (e.g. reduced RPM and stop the disk when unused for 30s).
That calls for a re evaluation of the situation.
Yev from Backblaze here -> we wrote about this in 2013 and have honestly not bought too many enterprise drives since, simply because they were more expensive and the benefit was negligible in our usage. Lately though Seagate has had a nice run of enterprise drives and they're well-priced, so we might be giving those more of a shot in the coming months!
I think having the drives spin constantly is the best for lifespan, because there are less spinups and thus components wear less.
So barring any manufacturing defects that I'm soure would be abundantly clear early on, it's down to logic or electrical failures.
I believe when you're buying as much drives as they do and produce the kind of reports they do they measure and compare before reaching that conclusion.
From what I can tell, they fixed those problems and went back to the high quality but slightly more expensive drives of old. The most interesting part is that Backblaze still buys mostly Seagates, simply because they are easier to buy in bulk and have a better price. Even though they fail more often the failure rate isn't bad enough to be a problem. Also, Seagate's failure rate has been decreasing as they release new drives.
To my knowledge, there have been innovations in the spinning rust (ahem hard drive) market since 2012. So why is HGST, a WD subsidiary, still making much more reliable drives?
Different design teams? HGST factories instead of WD?
[0] http://m.theregister.co.uk/2012/03/09/wd_closes_hgst_buy
That's more than my employer's Hadoop cluster...
(Context: http://archive.fortune.com/2006/11/30/magazines/fortune/obri... )
Even so it seems like a lot.
I started with 12 x 4TB in one NAS and I filled it in a little over 2 years. I just recently built a second NAS with 20 x 6TB to last me hopefully for the next several years.
The other remaining drives unaccounted for are in two workstations for local storage.
Uncompressed 4k @ 24fps, at 10bit color comes in at 324MB/s, so just being a wedding videographer could be 1.1T per hour of video. Any given small project could eat up 15-20T per project.
> personal use
Who needs that much storage for personal use?
1: Not sure what encodig
2: Napkin calculation, corrections welcome
Hint, hint, to clever founders: hard drives in Hollywood.
NAS #1 is 12 x 4TB with 2 drives used for parity.
NAS #2 is 20 x 6TB with 4 drives used for parity.
That takes away roughly 32TB of usable space under zfs. NAS #1 has 40TB usable. NAS #2 has 96TB usable. Total 136TB usable.
The remainder of the drives not accounted in the list above are in workstations for local storage and other random tasks.
I'd record to ProRes or DNxHR if I could, but there's no viable method for doing that on Windows without using an FFMPEG pipeline (which won't do 4:4:4 color sampling with those codecs and simply uses too much CPU at 4:2:2).
That's "home delivery" video, not the stuff you work on which would be far less compressed or even uncompressed, and would include reams of data thrown out entirely at the final production stages.
Think about the difference between "work" (mixing/mastering/production) audio data (uncompressed and at 24 or 32 bits) versus "consumer" audio (mp3/aac at 48/16).
I've bought 10TB of storage space every year for the past 3 years and don't see myself stopping. I need the storage space! If these data-heavy hobbies were instead my daytime job I could easily see myself needing 100TB+ in storage (and then 100TB+ of backups). I fill roughly 8TB/year in storage space (4TB of data + 4TB backup data).
Some people believe in redundant backups. I only have redundant backups of extremely important information - which are mostly text documents so don't take up much space. But if someone was storing 3 copies of all of their media assets you end up using a lot of space. For example, 10TB of video turns into 30TB of video. And 180TB is really only 60TB of data, which isn't that much for data heavy hobbies.
I returned it.
These tell us little about how good these drives hold up, it tells us more of the churn cost that backblaze has.
Edit: clarification
I, too, would like to see more regular SSD death matches.
Though of course you can fit a 12v/2a adapter in the backpack next to the drive.
I've briefly tried looking for this in the past before but was never able to find a well maintained source.
Sometimes. We do take it in to account but if a hard drive that costs less but fails more comes around and the failure rate in our environment is within some "tolerance" we'll get that drive, even if we might see some more failure. So we do use the data to inform our purchasing department of the "future costs" of drives, but a good deal's a good deal :)
I'd be very interested in generating raw data against the disks I maintain as well, might also be able to share it, which might be interesting since its primarily SSDs.
pods are 60 drives per unit now.
Is it Western Digital? Is it Toshiba? someone else? (afaik the HGST 2.5 and 3.5 production lines were sold to different companies -right?)
We also classify a drive as failed when it throws ANY SMART error, then it's diagnosed with Seagate's internal tool (which I believe they are open sourcing, if they haven't already) and put back into production
Seagate has given us good warranties on these drives as well. Yes a higher immediate failure rate (first 1 month) with the 8TB drives was annoying, but they more than made it right replacing the drives.
Never again.
I wish backblaze would share some data around accelerated lifecycle testing.
My photos are a) still on my sd card until it's full, b) on my raidz1 and raidz2 drives on my nas, c) backed up to flickr, free 1TB storage using https://github.com/richq/folders2flickr
Having only 1 copy of something that's easy to perfectly replicate is very silly.
Portable means the storage has a compact form factor and a case so you can carry it around (turned off). I have never seen anyone actually use a portable disk while moving. Everyone puts them on their desk next to the laptop. The only thing I can imagine is working on a train.
Good thing SSDs exist now.
In all seriousness though, I use backblaze and while I haven't had any HDD failures I have gone back to check on the backup from time to time and it's always looking good.