Backblaze Hard Drive Stats Q2 2019
backblaze.com
backblaze.com
So we started going on the theory that there were some marginal parts of the disk during manufacture, and exercising them would allow them to be remapped to good parts of the disk. Once we started including "badblocks -svw -p 5" in the burnin procedure, our drive failure rates dropped to basically 0.
I wonder if backblaze does anything similar to provision drives.
Recently I replaced some drives in our old Dell R720 servers (5 years old). They were showing media errors over several months but never went degraded. During the replacement, each server had another drive fail. So I ran badblocks on the RAID arrays over the weekend, just to make sure any other marginal drives or sectors had an opportunity to remap. No other drive failures.
Aside: I didn't lose any data, the arrays never went FAIL, but also I had migrated all our services off these machines. Love virtualization!
The bigger standardizing issue was the system. We use Supermicro barebones, and between memory and CPU changes, we would have to qualify a couple-three new chassis every couple of years. And those chassis were much more expensive than a disk drive.
Mostly, Supermicro quality was very good. With the exception of around a year or 18 months where the power supplies they were using all failed roughly a year after we bought them. We had older and newer servers with no problem, but around 30 in the middle all failed.
EDIT: the RAID6 array is not on macOS, but Linux. I have a TB2 external enclosure that allows me to easily insert drives on my mac. Hopefully the macOS version is as good as the Linux version.
> I wonder if Backblaze does anything similar to provision drives.
We do different things for different drives, but ALWAYS have a "burn in" period for new vaults (group of 20 computers that files are Reed-Solomon encoded across) that come online.
When I say "different things", when we deployed a vault full of recent Toshiba (?) drives it came up "too slow" and we figured out the OS block size and drive block size wasn't lined up "by default". It "works" but essentially a single write required reading two blocks and then writing two blocks. (sigh) I always wonder how many mis-configured PCs exist in the world with little problems like this quietly slowing down some user who never figures it out.
All of them.
From our IT guy who figured out the problem......
"If it's 4K native, then a multiple of 4K is a good starting point. If it's not then a multiple of 512 is ok. If it's SMR based then it's more complicated and depends on whether it's host managed or drive managed and the number of conventional PMR zones. The second part of the equation is alignment. If there's a partition table or logical volume mapping, then the partition or logical volume's offset needs to be a multiple of the native geometry otherwise the block size doesn't really matter because everything will be misaligned but it might be possible to use math to figure out a block size that cancels out the misalignment."
Sooooo..... consumers are hopelessly screwed and better hope Dell or Apple got it correct out of the factory.
More info from other IT guys.......
"These days I think parted gets it right most of the time, at least for individual spinning drives. SSDs are a bit weird because they tend to not want to share the erase block size, but I think the recent-gen ones tend to have enough firmware magic that it's OK. ... parted also has a flag for alignment type (--align optimal) which is supposed to choose an optimal multiple of physical block size."
I may use 12MB going forward, myself (rather than the "1mb covers all ur bases" I cited in my upstream comment).
2^20 bytes lines up with any known disk sectoring at the expense of "wasting" 1 mb, which wrt today's capacities is trivial.
If you want to declare a smaller offset, you can, but absent an explicit value, one (binary) mb is what you get.
Definitely not the sort of problem I expected to encounter during that project, though, so I can totally see how even you guys would hit stuff like that.
We started that burn in process when we could get "-p5" run in a day or two, maybe 120GB drives... Towards the end we were reducing it down to "-p1" and that took a week.
My 5 year old Dell R720s were running more like 2GB/sec, but that was across a RAID-10 array of 16x 10K drives.
I found that a bit ridiculous. My internet was so bad that I had to spin up an EC2 nano just to make a few billion requests and have them succeed. Took hours.
If backblaze wants to charge you on the way out (I wonder how much all those requests cost me), okay fine. But please just make a "delete bucket" button that does all this for us. If I was in a worse mood, I would've considered letting the account accrue balance until they did it for me.
Good to know I wasn't stupid. Couldn't shake the feeling at the time that I was missing something.
Agree it's stupid.
edit: actually the "don't close your browser window" warning leads me to believe that's just a client-side JavaScript solution which doesn't work with massive buckets (we had one with millions of objects)
[1] - https://docs.aws.amazon.com/cli/latest/reference/s3api/delet...
Edit: Found aws s3 rb --force, but this just automates the process of listing/deleting objects then deleting the bucket. When you have 7 million objects in a cloudtrail bucket, that can take days.
It seems that calling b2_delete_key is free. It's just not conveniently abstracted. Listing the items however if you don't already have a catalog is $4/M items
That doesn't seem correct. Each call to b2_list_keys can return up to 10 000 results, so it's $0.40/M items.
I guess this part from one of their answers fits your complaint quite well: But here is the thing -> YOU can work around this problem, the naive customers CANNOT. Honestly, they are too computer-illiterate. But even computer illiterate people deserve to have their files backed up, and they are the target market for Backblaze Personal Backup. [1]
From this point of view, I guess a "delete" button could be fatal for some of the older folks out there. But I don't use Backblaze, I don't know if my comparison makes sense and we both talk about the same thing. The Reddit quotes refers to their "Backblaze Personal Backup".
[0] https://www.reddit.com/r/IAmA/comments/b6lbew/were_the_backb... [1] https://www.reddit.com/r/IAmA/comments/b6lbew/were_the_backb...
I’m voicing this here because I was surprised at the UX. I was happy with my backblaze experience otherwise. My bill was incredibly cheap for the amount of data and access I was using. I’d consider using it in the future for things we generally default to S3 for.
The best way to deal with 'are you really really sure' is something like send an email with a link you have to visit as confirmation as that breaks you mentally out of the 'keep confirming without thinking' cycle which we are all at least slightly prone to given the number of confirmation dialogues we all see these days.
Is that ageism?
There is zero benefit to over-generalizing that I can see, and it is very likely to lead to errors.
Or do you not have personal experience working with the very elderly? The very elderly engineers are more into electronics (physical) and can run rings around younger folks on older electronics (repair / vacuum tube testing etc). But I've not seen high levels of interest in things like Backblaze personal backup API's.
People of all ages vary in their IT abilities, and IT abilities are obviously what matters in dealing with issues like backup.
There's just no reason to bring up age at all.
> How do those customers ever delete anything? Click every single file?
The current recommended work around is to write a very short "life cycle rule" in the web GUI that deletes any/all files 1 day after it is uploaded and let it run for 24 hours and come back to an empty bucket. Yes, I know this is lame. https://www.backblaze.com/b2/docs/lifecycle_rules.html
Just to be clear, there are two separate product lines at Backblaze: 1) Backblaze Personal Backup which is COMPLETELY automatic and you don't need to delete anything, ever, it is all automatic, and 2) Backblaze "B2" which is a toolkit for IT people and programmers. This is only a problem for the "B2" side, and the IT people and programmers (for now) can write a script or write the "life cycle rule". But we will get this fixed, it is on the roadmap.
> YOU can work around this problem, the naive customers CANNOT.
That was specifically in response to a Backblaze Personal Backup feature request.
> I guess a "delete" button could be fatal for some of the older folks out there.
Just to be clear, Backblaze offers two product lines: 1) Backblaze Personal Backup which is designed to be simple and easy to use, and 2) "Backblaze B2" which is designed as a toolkit for programmers and IT people. So in this particular case, we SHOULD have a "delete all files" button (for IT people), and we know this currently sucks (it's on a roadmap to add the feature).
You can select "some group of files" (like one folder) and delete them all at once from the web GUI, but if you have like 1 million files (which is a perfectly reasonable and normal backup) the web GUI will "time out" and fail to delete all the files in one operation.
The current recommended work around is to write a very short "life cycle rule" in the web GUI that deletes any/all files 1 day after it is uploaded and let it run for 24 hours and come back to an empty bucket. Yes, I know this is lame. https://www.backblaze.com/b2/docs/lifecycle_rules.html
To paraphrase my quote about computer naive users, the reason we don't prioritize this particular feature higher is the IT people (like yourself) are capable of doing the crazy work arounds like spinning up an EC2 instance. :-) But we will get to it, I promise!
In our defense, we have 5 open programmer job recs, and 5 open IT job recs, and we're having trouble finding enough (qualified) people who want to come help us build these features for customers. If you want to work in a really fun employee owned business in San Mateo, California where you can bring your dog to work, please come and join us! https://www.backblaze.com/company/jobs.html
There are a lot of vendors selling “refurbished” HGST drives. Do not buy these. As I understand it, there is no such thing as a refurbished hard drive (not economical). They just zero out the SMART data and sell it as new/refurbished. They even look new, but they’re not.
I made the mistake of buying one. It had a ton of vibration and started reporting bad sectors immediately.
If you want one of these drives, you’re better off buying an actual used one off eBay or something.
https://www.blizzarddr.com/recertified-hard-disk-drives-refu...
It simply makes no sense to actually open a US$ 100 (when new) disk drive to repair it (replacing parts that even in the factory cost money for the cataloguing and storage) to re-sell it at US$ 30 or 40.
If you go to the WD page mentioned in the article you linked:
https://www.wd.com/en-gb/products/wd-recertified.html
you will see how they are all "external disks in a case", those might be subject to a number of "returns" for reasons different from the actual disk inside (cosmetic damage, power adapter (if any), connectors, network card or wi-fi, etc.), so those may (while actually being defective as a whole) well have not a defective disk inside.
On the other hand, if they are so good, why would the warranty be limited to 6 months (or less).
Blackblaze uses f00{0..9}.backblazeb2.com which represents different data centers.
Currently f003.backblazeb2.com is live and returns a IP based out of the Netherlands.
ETA...sooner than you may think, but not Today!
Does anyone else use these posts to pick their own personal disks? If so do you see the data given as a good indicator of lifespan or is the home environment vs DC just too different?
it's shouldn't hurt to use their numbers to guide your decision it is no guarantee though.
Thanks!
For example, their 2% failure rate of Vendor X drives doesn't mean anything to you if you don't have enough drives. Your few drives might be in the 2% or the 98%. The 2% failure rate is only applicable to the population of drives as a total. It doesn't tell you anything about your drives in particular.
Now, on top of that, their usage patterns don't necessarily reflect what a typical user would see on their personal computer. I suspect that their read/write patterns skew towards the write-once, read-rarely category. Even for a typical enterprise datacenter, their usage patterns are going to be different.
If Vendor A has a 2 percent failure rate and Vendor B has a 3 percent failure rate it’s still valid to conclude that you would be less likely to have a failure with Vendor A even if you only ever buy one drive.
Home use being not comparable to their usage patterns is a valid argument, but your statistical argument seems like nonsense to me.
It’s only when your sample sizes get sufficiently large that these failure rates start to have a good predictive value.
Living in the future is cool.
- warranty and customer service - price - noise - performance
You can get lucky and have a drive that lasts 10 years, or you can get unlucky and have a drive die in days. Either way, you'll need a solid backup strategy and be able to be without the drive for a week or two.
That being said, I appreciate the data that they put out, and it makes me want to use their service more. I think it's interesting that they continue to use Seagate drives more than most others despite them having higher failure rates, which reminds me that you really need to look at it from a cost/benefit perspective. It's also interesting to see just how reliable drives are (on average) even when used improperly.
In short, no, I don't think they should be used when deciding which drive to buy. All drives are pretty good, so find the one that's going to make ownership and replacement the most pleasant.
> The important things for me when considering a drive are ... noise
For a drive in the same room as me, I TOTALLY agree and I'm surprised this isn't mentioned more often. Before I switched to all SSDs in my desktop computers I ran cables through the drywall in my apartment so I could put the (noisy) computer in a closet in the other room and put the monitor, keyboard, and mouse in the quiet room with myself.
Besides the wonderful performance, SSDs are magical to me because they are quiet. I am very willing to pay a premium just to be free of that drive seeking and grinding.
For personal drives in small quantities, you're basically operating on luck. You can get a bad HGST drive just like you can get a bad Seagate drive. I've had drives from HGST, WD, and Seagate and all of them have lasted far beyond the warranty expiration.
When I look into drives for small-scale personal use (NAS, desktop, etc), I check recent reviews about warranty and customer service, then buy based on price and features. I also buy drives based on the use case (NAS drives for NAS, desktop drives for desktop, etc).
You need to add "cost to repair" into your equation. At my old job we used to figure $300 to replace a drive (scheduling maintenance window with customer, taking backup, migrating/quiescing services), replacing drive, verifying services were happy, wiping or destroying old drive, managing RMA, qualifying/burning in new drive).
YMMV depending on policies and procedures. If I were Backblaze, I'd design the system to work equally well with 40 drives as with 45, and consider just leaving bad drives in until a handful of them need replacing in a chassis. And the replacement be basically automatic.
This version of the report is much closer than they've been in the past. Seems like Seagate might be making some improvements. IIRC, it was not uncommon for Seagates to be a 10x higher failure than HSGT in the past.
They designed their own replication scheme which allows quite a bit of disk to fail. Considering that, using RAID would be absurd so I'm pretty sure they are all configured as JBOD, thus allowing this kind of replacement.
While they do have multi-failure redundancy (17 data, 3 parity), it wouldn't be a good idea to push the limits regularly because that means you're risking the loss of actual customer data. And adding more parity to give you extra wiggle-room would result in both storage efficiency losses and performance losses. It probably wouldn't be advisable to expand the "stripe" size to cover an entire vault (to increase the independent disk failure tolerance without reducing the data-to-parity ratio) because you'd kill performance (the matrix multiplications necessary to do Reed-Solomon have polynomial complexity).
[1]: https://www.backblaze.com/blog/vault-cloud-storage-architect... [2]: https://www.backblaze.com/blog/reed-solomon/
> You need to add "cost to repair" into your equation. At my old job we used to figure $300 to replace a drive
This is TOTALLY true, and Backblaze does add in the cost to repair. At our scale, we have full time datacenter technicians that are onsite 7 days a week for 12 hours a day (overlapping shifts) so that we can replace drives and repair equipment within a certain number of hours.
Every single day our monitoring systems kick out a list of about 10 drives that need to be replaced, and the data center techs try to make sure there are zero failed drives when their shift ends. (We leave failed drives alone for up to 12 hours at night as long as it is just one failed drive in a redundant group of 20.)
But if you are operating with FEWER than our 110,656 drives then you can't have full time repair people standing around paying their salary. So you are really going to need to pay much more (per drive) than we do.
One of the absolutely amazing things about "cloud services" like Amazon S3 or Backblaze B2 is that (hopefully) we can actually operate it at a lower cost and higher durability than you can operate yourself. That may not be true at either end of the spectrum - if you have fewer than 10 drives it may be cheaper for you to operate it yourself, and if you have more than 1,200 drives (the deployment size of one of our vaults) you might want to consider cutting out the middle-man (Backblaze) and saving yourself money. But I make the bold claim we can save you money in that sweet middle ground.
One of the factors to consider is the cost of electricity in your area. Backblaze pays about 9 cents/kWh which is "pretty good". Up in Oregon you can get deals for 2 cents/kWh and beat our operating costs, but in Hawaii you are at 50 cents/kWh and unless you have solar panels you pretty much better host data on the mainland. :-) I think TONS of people (mistakenly) think that after they purchase a hard drive, operating it is "free". Electricity is one of Backblaze's major costs in providing the service. Our electrical bill is more than $1 million per year right now.
Some people think we (Backblaze) get "magical good deals" by purchasing drives in bulk, and it isn't as good as you might think. Sure, we get some bulk discounts, but think 5% or 10% better than your retail price for 1 unit. And you can probably get THE SAME DEAL as Backblaze if you are purchasing 1,000 drives in one purchase order.
Backblaze is a pretty good deal. I'm currently trying to figure out if I should take my video files that I'm editing for my personal movies and store them on Backblaze (transmission time being the issue there) vs. getting a 10TB drive to put them on or a ZFS array, vs. just deleting the source. The latter has a very competitive price. :-D
In some environments when a drive goes down, everyone goes into "red alert" mode and they have to rush to repair the drive or replace it. In our case we don't have to drop everything and run to it - so it helps keep the overall replacement costs down.
And reality sucks: right yesterday I lost one hdd from Seagate and the second one is making noises. Ohh boy
I think I saw double-digit failure rates from certain Seagate models a few years ago, though :(
> No wonder Backblaze has replaced most of their inventory with these two brands.
We Reed-Solomon encode each file across multiple drives in multiple machines, and we always use "enough parity" that the failure rate won't result in data loss.
So for us, we have a (pretty simple) little spreadsheet that takes into account drive failure rate as a cost, along with drive density (more dense drives means renting less data center space) and we let the spreadsheet kick out which drives to buy based on the cheapest total cost of ownership. Honestly, we're not brand loyal AT ALL, and we're not afraid of a higher failure rate (other than that raises the cost because we have to buy replacements for failed drives).
With that said, if an individual purchases ONE DRIVE they might value a lower drive failure rate differently than Backblaze does. But honestly, if there is a 1% drive failure rate per year or a 10% drive failure rate in a year, you should still backup the data so you can sleep at night.
We recently replaced all the 15K Seagate drives in our dev/stg 6 year old server with Samsung 860 EVOs, because I was replacing 1-2 a month (out of ~20 disks).
I've also had good luck with Toshiba disks, though I have a much smaller sample size.
I could deal with their quality issues if their RMA process was at least reasonable. With WDC it's a credit card number (if you want an advance RMA) and a web form to request an RMA. With Seagate it's a game to try and buy "official" packaging, a tooth-and-nail fight for an RMA number and a long wait.
They actually refused to RMA the spindle-locked drive because the PCB was damaged (well, no duh). Thankfully the supplier was more lenient.
My home system using WD Reds has been great. It is larger than a normal home storage box (24 drives), and it gets substantially more use than normal (backing store for a bunch of VMs, mostly). I've replaced one drive in five and a half years.
Contrast with storage servers built for work with WD Golds. Those are 45 drives each and used for database backup workloads. We replace about a drive a quarter (of the 90 total), and this has been pretty consistent for a bit over two years.
Not sure if I got great Reds or if Golds are just terrible, but that's what I've seen.
Honestly, all hard drives are pretty solid these days, and I honestly don't think one manufacturer really stands out enough that individuals can really expect a given drive to last longer than another. Statistics just don't work on a sample size that small. You may get lucky, or you may get unlucky, it really just depends on which drive you happen to get.
I make my personal HD buying decisions based on:
- warranty and customer service - price - noise - performance
Both Seagate and WD have about the same warranty, and they both seem to be pretty good with customer service (check recent reviews though, this can change). If you consider them equal in that department, pick based on the other three metrics.
I was about to buy some Seagate drives for a NAS, but ended up with WD because they were on sale. I was a little worried about the noise from Seagate, but the price and performance won me over, but I switched at the last minute because the price changed.
Now, if you're going to build a data center with a large number of consumer drives, absolutely use Backblaze's results. If your just going to but a few, recognise that all drives the test are reasonably reliable, so whether your drive dies early is mostly luck (assuming you treat them nicely).
They have the data spread out across enough drives with enough redundancy to live with it.
That being said, I had a pair of 1 TB Seagates in RAID 1 that had over 7 years of runtime, which I retired for a capacity upgrade.
At first I thought Western Digital is retiring their WD brand for HDD, and focus on SSD, but that is not true I find no information on it. There are still Red, Blue, and others. As a matter of fact, it is the HGST brand that is getting killed, but the UltraStar range lives on.
So I am assuming the article is saying, the specific breed of HDD, Western Digital used to provided is no longer available, and it will be replaced by UltraStar, which used to be a HGST brand but are now under Western Digital ?
What happened to 20TB HDD which Seagate promised for 2020. Are we no where close to it? And MAMR Drive? Has Price / GB continued to decline, or given the situation of HDD market, the HDD maker are milking it? ( And I don't blame them )
A batch of drives from the manufacturer had some sort of firmware issue. It affected a lot of users.
Essentially their data was put into read only mode by our service. Two weeks went by before the manufacturer came out and worked on replacing the drives and migrating the data over for us.
Unfortunately the data center admin at the time didn’t have backup servers and drives in place and lost his job over the snafu once the drives were replaced.
Always have backups of your data and make sure the drives are good quality.
Most home users don't need this, it's cheaper and better to sync to a 2nd drive nightly. This gives you a local backup and redundancy, because you can swap to that drive if your first drive fails. If you're doing RAID, you still need a local and offsite backup.
I used to do this with windows. Yeah, if you have corruption, you copy it, but so does a raid? At least this way your safe from fat finger errors and you have a quick access backup.
With 4TB and the price being $0.005/GB per month[0] that's $20 per month.