Backblaze hard drive reliability stats for Q3 2016
backblaze.com
backblaze.com
Guess you are not throwing those away so what are you doing with them.
https://www.backblaze.com/blog/hard-drive-reliability-q4-201...
No one really expects that rotational rust will get much faster, and in fact history shows that, compared to the increase in density, the increase in transfer rates are laughable at best. Between 1990 and today you are probably looking at a 20 000 times increase in density, yet transfer rates only increased around a factor of around 150-200. [In fact, from the early 1960 to today it's only a factor of about 1000. There are quite possibly few performance metrics that increased so slowly as disk transfer rate].
Why is that?
Increasing density does only marginally increase transfer speed: Most density increases are achieved by packing more tracks onto the platter, while storing more sectors per track plays a minor role. But a single R/W head can only read a single track, not parallel tracks, hence speed only increases if you pack more sectors into each track, not by increasing the number of tracks. That's why between today and 10 years ago performance in desktop or server drives only differs a little, compared to the capacity increase to 8+ TB in a 3.5" drive.
More platters also don't help in transfer rate, because the alignment of all heads on the actuator is fixed: at any given time only one platter and one R/W head is used (locked to the track). [More platters can help reduce seek time in certain scenarios though]
More disks on the other hand...
You could probably do it if the platter diameter were reduced, which would then accordingly reduce capacity as well. At that point you could just as well use two drives and also get lower overall failure probability. Or you use SSD caches, or memory caches, ...
Perhaps another option is to go back to the 1980's hdd designs where the arm moves straight down the radius of the platter. This design might permit multiple heads on the same arm. I'm sure all this stuff has been researched thoroughly.
Either way, this doubles/triples the probability of mechanical failure
At my first job, sysadmin and programming on a PDP-11/44, 1-2 fascinating Winchester drives were procured for it. They had clear plastic covers, and you could see everything, the disk, the actuator, which had two heads, and it was by and large square, I'm almost positive it was rotary, not solenoid based straight stroke like many '70-80s drives, CDC's famous line in particular.
Even less than that. Tracks are 100-300 nanometers wide; the head flies about 3-6 nm above the platter.
The actuator mechanism literally locks onto the signal encoded in the magnetic track and follows it as the platter rotates underneath the R/W head.
Higher density means more data per track, not just more tracks per disk. You get an entire track per revolution so a track with more data is more MBps. So linear reads on a higher density drive are faster, and semi-linear accesses (ie, reading two files that are next to each other) do get faster.
I remember reading a story about a guy who built a drive array with high capacity 7200 RPM drives that got within 20% of the performance of the 10K RPM setup they had, by partitioning the drives at the same capacity as the 10K equivalent. The head only had half as many tracks to traverse, so worst case access time was better, and the higher density made up for the lower RPMs.
You parent comment is right though, there are only small changes in bit density on the track in recent years so the bandwidth is not improving by much.
You don't double the write throughout on a disk by doubling the number of tracks. You need more platters and/or sectors per track to do that.
I literally said that:
> Most density increases are achieved by packing more tracks onto the platter, while storing more sectors per track plays a minor role.
Instead of 50-100MB/s, you can get 4-8x the speed in large linear transfers, which helps get dead racks back up faster and would work quite well in backblaze's backup model.
Your block sizes will be huge, but I think in backblaze's case that doesn't really matter so much.
The smart move is to keep several smaller arrays instead of one big one. This lowers risk as well. I dont put in anything bigger than 7 disks into production. Past that I'm just asking for trouble. Its better to have 4 7 disk arrays than one 28 disk array. A drive fail means a quick rebuild and a restore is going to be 1/4 the time.
That takes ~11 hours to fill. At 400x the density it would take 20x as long or 9 days. I don't think HDD drives are hitting 400x the density any time soon but if they did it would be a problem.
However, in an array you could take a month to fill a drive to 75% without causing to much trouble. Assuming you had enough drives. That's around a ~80PB limit drive. IMO, the real issue is it would take another month to download all that data. Relegating HDD firmly into archival storage.
PS: I don't think rust is going to get into those density's making this far less of an issue.
For example when I search for HGST HMS5C4040ALE640 on Amazon I get a dealer selling old out of warrantee drives as new.
https://www.amazon.com/HGST-MegaScale-HMS5C4040ALE640-Coolsp...
I get similar results with many of the other drives listed and with other websites such as NewEgg.
https://www.backblaze.com/blog/hard-drive-reliability-update...
I don't read this as 'which drive to buy' but more as 'which drive not to buy'.
Disclaimer: I work at Backblaze. I know you didn't mean that as an absolute, but I just want to point out 100% of drives fail. It's my OCD that makes me point this out. We have NEVER found a drive that lasted forever. There are two types of drives: 1) those that have already failed, and 2) those that are about to fail. For any data you would be annoyed to lose, you need three copies in three locations with three separate vendors (three different pieces of software that don't share any lines of code).
Also, not sure there are three pieces of storage software that I trust and are readily available to me.
When you buy hard disks ONLY buy from Amazon or Newegg directly - never buy from a 3rd party seller on their site. Especially for hard disks there is too much fraud, and for a hard disk especially the risk of data loss makes it just too risky (unlike other items).
Agreed. A few times I've purchased third-party disks and found (via SMART data) that the drives were well used despite not being sold as such.
We found this, as well. And we stopped trying to get the HGST drives after we got a bad batch from a seller on Amazon.
Also seems there's no minification or combining of stylesheets/js and there are query strings on those static assets which is going to discourage caching.
No wonder you need a datacenter to handle that kind of resource punishment!
There are plenty of reasons to stick with Wordpress in a decent sized corporation but if not switching to a static site at least stick W3TC on there so you're minimising your server load and serving out static html and minified/combined resources.
You could then consider using Varnish in front of Apache or maybe nginx with a FastCGI cache.
I"m sure you've got some folks in the team who could whip up a W3TC install in 10 minutes.
Because if it is it from the team that currently can't keep a blog post online when you get a few thousand concurrent visitors, so you might keep yourself open to suggestions and perhaps undertake the BASIC best practices of keeping a Wordpress site up under load.
If nothing else it shows a basic lack of planning for what you know to be a massively popular post, so turn a little of that judgement back on yourselves.
It's possible easily handle tens of millions of hits a day on a tiny VPS if you do even some basics right[1] and that was without any particularly extensive optimisation.
[1] http://reviewsignal.com/blog/2014/06/25/40-million-hits-a-da...
EDIT: I may not be allowed to reply to the comment below due to HackerNews restrictions so incase the option doesn't become available in the next while I'll just say I accept the answer below gracefully, withdraw my daggers and take a calming beer at the end of a long day :-)
I'm wish you continued success and look forward to the next post.
*Edit -> to your above edit -> I think if you expand the comment by hitting the "time submitted" link you can leave a reply, thus subverting HN :P
Yoast is a culprit of performance though.
Don't forget the plugin query monitor and http2 doesn't need bundling resources ( I suppose)
Only asking 'cause it's my main data hard drive...
Among the least reliable drives I saw in previous reports were Seagate 3TB drives (supposedly they had worse reliability than the legendary IBM Deathstars) and after reading about how 3TB drives were designed across manufacturers years ago during the flooding crisis I decided to avoid 3TB drives entirely. Seems like my decision is finally getting some data to back it up now in hindsight (no pun intended).
Ok. But why ? Technical obstacles such as having to deal with distribution diversity - or is it a way of market segmentation ?
Let's be honest: when people see unlimited, most think "I don't have to worry about how much I'm storing" but a small group thinks "How can I take advantage of this?"
Not supporting a Linux client fixes that quickly.
Fixed that for you. We're just trying to maximize our resources and minimize cost. Nothing wrong with that.
I don't want to take advantage of it. I just happen to have 8TB of data to back up...
But backing up that much data over the internet isn't practical in any case.
Disclaimer: I work at Backblaze. The underlying base of original client backup software was originally written from scratch on three platforms simultaneously: 1) Windows, 2) Macintosh, and 3) Linux. It was designed that way from the beginning. This code continues to compile every time we do a client release, simply as part of the process. However, it is entirely lacking a GUI layer and an installer - those were never written. The underlying backup engine runs even when the user is logged out or the GUI has stopped working.
So it is technically possible, but along the way we released Backblaze B2 (storage API) which not only supports Linux, we assume Linux is the primary customer! We're seeing if that can satisfy the Linux community. Backblaze B2 is a large ongoing effort consuming a lot of our software developer's time.
A note about limited resources: Backblaze never really raised any funding, there are no deep pockets, so we can ONLY hire an additional programmer when the products we sell throw off enough money to pay that salary. We run on really tight margins (thus our obsession with failure rates of drives) which is fabulous for our customers, but not so great for hiring lots of extra help to do projects like a Linux GUI. :-)
Why would you want a CLI per hoster, if there are CLIs that target most hosters? Service?
Not so reliable I gather