Backblaze Hard Drive Stats Q2 2020
backblaze.com
backblaze.com
* A hard drive may go missing in the data from one day to the next without appearing as "failed" -- does this mean you've taken these hard drives offline for some reason?
* A hard drive may appear as failed one day, and then appear the next day with a failed=0 flag. Does this mean that these drives have been serviced and returned to the front lines?
Since you mentioned cost, do you have data on average price per GB, both acquisition and electricity cost / year? With how many drives you guys bought there should be enough datapoints for a pretty neat graph.
Edit: I use rclone as the backend for duplicity, so you can also chain it through another tool with different encryption and use rclone as just the transfer engine, getting all the benefits of rclone's providers with the benefits of duplicity's backup strategies.
If you're comfortable setting up a cron job, it's a great fit. I use it to back up a 1.5TB Samba directory and wind up paying about $5/m for B2.
You want tarsnap.[1]
Edit to add: Colin Percival is the author of scrypt[2] and has worked extensively with FreeBSD's portsnap, so he knows what he's doing.
To be more precise, I wrote FreeBSD's portsnap. (Also, freebsd-update.)
But one of the benefits of backblaze is the simplicity. Simplicity of setup, backups, and restores. If you muddle with that by encrypting before giving to backblaze you lose out on part of the value.
It also usually seems when people roll their own it opens up risk of forgetting something. BB is easy, set up their encryption and you should be fine.
Just some thoughts.
For restore, though, decryption happens at the server end. You have to supply your key to their server, which decrypts the data at their end, then sends you the subset you are interested in restoring.
See [1].
This is the only thing putting me off backblaze
That surely assumes an upper limit of the likelihood of a drive failure. There was a perception that the quality of 3.5" floppy disks declined drastically in the early 21st century. Must we not fear something similar for spinning-rust hard drives once most everyone uses SSDs?
Unless you have an uptime bug in the firmware where all your drives die at once:
* https://www.zdnet.com/article/hpe-tells-users-to-patch-ssds-...
At that point, if there's not contingency redundancy built in (See below), it's really a matter of how long it takes to replace a drive (in both identifying the problem, physically replacing the hardware, and replicating data to it). There's a lot of (fairly simple) math involved in running down those numbers, but based on the percentage of drives that fail in a quarter, I think it would take both a spectacular run of bad luck combined with negligence on their part in making sure redundancy levels are kept over a longer period to actually have problems.
> Is there not the danger that the quality drops drastically to the point that one would need an unreasonable number of copies?
I think the very simple way to look at this is that space capacity and automatic redundancy checking can account for a lot of bad drives. E.g. if a drive has 100 chunks of data all copied to 100-200 other drives and systems (such that there are three copies of any chunk), that the data exists three places, and if that drive dies and the system detects those 100 chunks are now only exist in two places, it can immediately locate 100 locations that have capacity to receive a chunk and start replicating data to keep the level of redundancy they need. Even if there was a very large set of bad drives, they would have to all go bad in a very short time frame, short enough that the couldn't be physically swapped out and data couldn't be copied across the network, for it to cause a problem.
At least that's how a system like this could be developed, and my understanding is that Backblaze's system works like this to some degree.
Typo of the year, "Boing 737MAX" sounds more like a basketball than something I'd want to fly in.
This big takeaway from these numbers is how dramatically low they are in comparison to drives 5/10/25 years ago. If you treat them reasonably, modern drives are rock stable.
Basically, divide up each disk into 1TB partitions, then RAID the 1TB partitions together. Synology uses this for their Hybrid RAID. Obviously care needs to be taken to not confuse partitions but this lets you use disks of differing sizes.
Unraid also offers RAID with differing sizes for disks, it has quite different failure modes though but it's been fairly reliable for me.
I'm using a combination of striping, RAID and mirroring depending on the content, using lvm and mdadm instead of ZFS for the increased flexibility. The server itself contains a HP R410i RAID card which drives 8 internal 2.5" SAS drives, the thing could be expanded for not that much money to a 16 x 2.5" array but I see no reason to do so given the higher prices for 2.5" drives and the abundance of capacity offered by the DS4243.
Not really, disks are cheap enough that I'd rather keep things simple. I have redundancy in various ways in addition to using Backblaze.
For my own use, I've found it isn't worthwhile expanding beyond eight drives rather than getting rid of smaller ones - adding additional ports costs money and it takes more power and space.
It might not be the cheapest, but 4+1 drives is more efficient than 3+1. Currently I have three 8TB drives: 1xWD Red and 2xIronwolfs (wolves?) with one running as parity (SHR). No complaints. It does everything I need it to do. The interface is simple and everything "just works" as anticipated. If/when I run out of space I plan on adding another 2x8TB drives, probably more ironwolves.
I wouldn't go bigger than 8TB as they already take DAYS to add to the array.
Migrated from a 1512. It took probably about 3-4 days to migrate everything. The new box handled it easily. The old one was struggling a bit with rsync. As the default is ssh with it. Once I had to downgraded the encryption then it went a lot faster (from 30MB to 90MB which was close enough to the limit of the cable to not worry).
This year I broke down and bought a Synology (918+). It's not perfect but I'm still a HAPPY happy camper --> I'm at a point in my life where RAID is a means/tool, not a goal / fun project to pour lifetime of hours upon ¯\_(ツ)_/¯
I have an old TS140 with an i3 that can take 4 HD, has ECC RAM with 4 slots. I'm currently exploring setting it up with Freenas (or Truenas after the merge) based on ZFS. So far it looks pretty good and performant.
For off the shelf NAS, Synology options look good..
I found that for my size (~10TB external), Backblaze went from very cheap personal to very expensive business pricing :-/
I previously had a HP N54L with FreeNAS, modded with extra two extra drives to get 6x4TB and that setup was very simple to do.
Now I moved over to a custom one which is a Fractal Node 804 case, i3 8100T, C246M-WU4 mobo and 32GB ECC DDR4 RAM. It has a SF450W Platinum PSU, and it pulls less than 40W in use - could get it lower with some underclocking. This runs Debian Buster with ZFS on Linux, and it's been rock stable for over half a year now - moving the drives over from my N54L was surprisingly simple. Plex transcoding works fine as well. Definitely recommend it.
I shopped around and waited for sales/price drops and got it all for under £500 (not including drives). I don't the pricing if your're not in the UK, so it may be cheaper elsewhere to do another combination.
As for drives wait for good deals on WD Elements/My Book Desktop External Hard Drive and shuck them. Note, only do this for +8TB drives (these are CMR drives atm) as under that size the drives are SMR which you should avoid like the plague for RAID/ZFS as they have been reported to fail if you ever need to rebuild. So then you're looking at 4x£120 which is another ~£500ish investment.
If you want to save money get a preowned box like the N54L off eBay at ~£100, and spend the money on drives, then stick FreeNas on it and if you're only file-serving the box is good enough.
I think the custom box is a bit overkill but it's got so much room for upgrades, I plan on getting 8x10TB drives in there one day, plus you can run VMs on it as the i3 processor isn't too bad.
Synology, and you should also look for and ONLY buy models which support Btrfs. Silent file corruption is very real.
Personally I am waiting for Qnap's Hero Edition OS, because I just love ZFS.
Other than these two, I am not aware of any other consumer NAS which support either of these file system.
Or may be someday Blackblaze decide to enter NAS market :D
It would also be interesting to see the time-dependence (i.e. does AFR really look U-shaped over the lifetime of a drive?). That would require a dataset with every drive used, along with (1) number of active drive days, and (2) a flag to indicate if the drive has failed and of course (3) which kind of drive it is. Does Backblaze offer this level of granularity?
EDIT: They offer the raw data dumps!
https://www.backblaze.com/b2/hard-drive-test-data.html#downl...
Backblaze, god bless you.
I've always really enjoyed reading Backblaze's reports though and have made past buying decisions based on their information.
I wonder if it might be possible for a community sourced version of these reports? A small app that checks SMART data and sends it to a central repository for displaying stats? Many from the self-hosted/homelab crowd are running these shucked drives so there has to be a large pool of stats out there if it can be gathered?
Incidentally I've tracked down a HDD (spinning rust) performance (and later breakage) problem to a very loud (and I mean awfully painful, intense vibration on your eardrums) alarm test every wednesday...
Do you monitor the noise in your DCs? Maybe for high peaks?
HGST drives win the prize for predictability. Boring, yes, but a good kind of boring.
Edit: See comment below, looks like this happened a while ago, so not likely a real factor :)
Explains why HGST brand drives are harder to come by I think.
That can make your data look corrupt. Yet reformatting (or just putting the drive in with the old motherboard) will magically fix them.
[0] https://www.backblaze.com/blog/wp-content/uploads/2020/08/Ch...
Years later people noticed that the drives were highly reliable again and the Deathstar problem was a fluke. But now it's too late. All of the HDD manufacturers have been gobbled up by just two companies and all of the interesting competition is on the SSD side.
Also, if you don't use one, please elaborate if you can why the obvious MTBF delta projection idea is bunk. I'm sure your talented team has thoroughly thought through all of this
It would be interesting if these reports included prices, but that might be a problem for Backblaze to reveal so much about their business operational costs.
They had a disassembly line gutting external hard drive enclosures for the hard drives, too. One of the drives in my array was acquired via a similar process.
You should read the whole story. I should probably read it again myself.
For home users a dead or dying drive is such a headache that it's usually worth spending the extra bucks on something that is less likely to fail on you.
Long ago we hit a point where power and cooling, not floor space, were the biggest bottlenecks for many data centers. Those numbers shift around, but during that era a bunch of people noticed that it's less labor intensive to fill an entire rack at a time. Put a few extra racks in as spares, and as hardware fails you power them off and power up new machines in the spare rack. When X servers in a rack are dead, you migrate everything else out of the rack, then strip it to the rails and put brand new hardware in it. You increase your server density twice in that case.
Event-driven maintenance is disruptive to other work. It tends to have a ticking clock involved, so you likely assign it to more experienced people, which is 2 opportunity costs in one. Rack-at-a-time is batch processing. You can schedule it, you can assign multiple people to it, and speed is less of an issue. I might have the new guys stripping a decommed rack, use building a new one as training for four people and one teacher. You put a third of it together, I'll come tell you why the wiring is wrong, lather, rinse, repeat.
Lets say you found that HPE SSDs were super reliable over your test period so you decided to put everything on a bunch of new drives.
Then this hits, and 100% of your drives fail at the same time: https://www.pcmag.com/news/time-to-patch-hpe-ssds-will-fail-...
While these stats are a great analysis of random failures -- that's only one type of risk.
The worst thing you can do is to put all of your eggs into one basket. And bulk ordering might get you a set of drives all from the same batch. Bad batches will tend to all fail for a similar reason. Drives fail on a bell curve, right? So the first drive may fail way before any others, but in a RAID array, rebuilds are stressful. Eventually you will hit a statistical cluster. Multiple drives failing close together. If that happens during a rebuild, you will lose a RAID 5 array. If you are very lucky, your RAID 10 array loses two drives in the same mirror. If you have two failures during a long rebuild, even RAID 6 won't save you.
I just bought a Synology box for home. This is my third and probably final RAID enclosure for personal use. I was having trouble finding Backblaze-tested drive models to populate it, so I filled it with drives from a Drobo and kept looking. Initially I had populated the Drobo with 4 drives I bought at once. When one failed, I bought 2 HGST drives and replaced a pair. When the new drives arrived, I started trying to cycle them through, and one of the drobo drives failed. I'll give you two guesses which one.
There is, as far as I can tell, no prosumer multi-disk filesystem that uses consistent hashing to stripe+mirror files across an arbitrary number of disks, instead of the heavy linear algebra RAID5 relies on. It requires touching the whole disk on every rebuild, and I believe that's why Object Storage is slowly taking over from the top end. It's a simpler form of redundancy.
I hope that it's worked its way down to my price range by the time the motherboard on the Synology burns out.
There is a failure mode in disks which can be modelled as something fails on the drive, but it continues to work fine for maybe months or years until the next power cycle, upon which it then won't work.
Obviously that's a problem for redundancy schemes because you think you have plenty of redundancy till there is a power outage and suddenly loads of drives fail at once.
I have never seen any of your reports measuring or reporting on these 'fail after power cycle' events, which is surprising.
From the behavior of my RAID (which also uses Reed Solomon, doesn't it?) it feels like repairing an array takes time proportional to the size of the drive, not the size of the contents, and it feels like a waste. But it's possible that my comfort levels for available disk space are a lot more conservative than other people's, and so the difference is less pronounced in a 'normal' storage situation.
For instance, an array that's at 80% capacity takes 25% longer to rebuild than I wish it would, whereas an array that's at 66% capacity takes 50% longer.
This is only true for drive-level raid rather than filesystem level raid, or a non-raid solution like ceph's replication.
ZFS's filesystem raid can repair a raid in time proportional to the amount of data stored in it.
mdadm and raid controllers aren't aware of which parts of the block device are in use or not, and thus have to repair the whole drive.
It's exceedingly likely that backblaze's solution does not require repairing entire block devices, but rather is likely to be closer to ceph, where only the in-use portion of a failed drive must be considered / must find a new home.
I think raid and distributed storage systems (like backblaze or ceph) are more different than they are alike.
> From the behavior of my RAID (which also uses Reed Solomon, doesn't it?)
Maybe. mdadm raid5 doesn't, nor does mdadm raid1 or raid10. I think mdadm's raid6 does.
You could then for example have 18/2, and then group together 400 drives in a 2nd layer of 19/1. Hey, I reckon you could do 19/1,39/1, reducing your storage costs 7.5%...
Sure, the worst case rebuild cost is much worse, but overall data loss probability is far lower, and a 2nd layer rebuild is a very rare event, and in that case, your customers totally prefer a few extra seconds latency over an email reporting their data is lost...
I assume you mostly do streaming rather than random writes, so the overhead is evenly spread amongst the disks, and is the same 15% as your current scheme.
https://en.wikipedia.org/wiki/Bathtub_curve
It'd be interesting to see if BackBlaze sees this in their drive populations.
Also, why can't I find HGST drives for decent prices? On Amazon they are either refurbished drives or crazy expensive prices for new ones?
Just a heads up: there's no such thing as a refurbished hard drive. It's not economical. Instead, those drives are actually just used drives with the SMART counters cleared.
I've been personally burned by this.
That being said, there are definitely some sketchy drive resellers on marketplaces like Amazon who just clear the smart data on old drives and sell them as refurbished, or even new.
Because they're discontinued, Western Digital rolled them into their server line, and people like me are snapping the remaining stock incase drives in our RAID die.
Is this underlying manufacturing error tolerances, or is this shipping/deployment effects, or is this .. Aliens?
Just looking at the pricing pages, pure storage is cheaper at B2, but the API calls and egress are not free.