Does anyone else use these posts to pick their own personal disks? If so do you see the data given as a good indicator of lifespan or is the home environment vs DC just too different?
Does anyone else use these posts to pick their own personal disks? If so do you see the data given as a good indicator of lifespan or is the home environment vs DC just too different?
For example, their 2% failure rate of Vendor X drives doesn't mean anything to you if you don't have enough drives. Your few drives might be in the 2% or the 98%. The 2% failure rate is only applicable to the population of drives as a total. It doesn't tell you anything about your drives in particular.
Now, on top of that, their usage patterns don't necessarily reflect what a typical user would see on their personal computer. I suspect that their read/write patterns skew towards the write-once, read-rarely category. Even for a typical enterprise datacenter, their usage patterns are going to be different.
If Vendor A has a 2 percent failure rate and Vendor B has a 3 percent failure rate it’s still valid to conclude that you would be less likely to have a failure with Vendor A even if you only ever buy one drive.
Home use being not comparable to their usage patterns is a valid argument, but your statistical argument seems like nonsense to me.
It’s only when your sample sizes get sufficiently large that these failure rates start to have a good predictive value.
Thanks!
- warranty and customer service - price - noise - performance
You can get lucky and have a drive that lasts 10 years, or you can get unlucky and have a drive die in days. Either way, you'll need a solid backup strategy and be able to be without the drive for a week or two.
That being said, I appreciate the data that they put out, and it makes me want to use their service more. I think it's interesting that they continue to use Seagate drives more than most others despite them having higher failure rates, which reminds me that you really need to look at it from a cost/benefit perspective. It's also interesting to see just how reliable drives are (on average) even when used improperly.
In short, no, I don't think they should be used when deciding which drive to buy. All drives are pretty good, so find the one that's going to make ownership and replacement the most pleasant.
> The important things for me when considering a drive are ... noise
For a drive in the same room as me, I TOTALLY agree and I'm surprised this isn't mentioned more often. Before I switched to all SSDs in my desktop computers I ran cables through the drywall in my apartment so I could put the (noisy) computer in a closet in the other room and put the monitor, keyboard, and mouse in the quiet room with myself.
Besides the wonderful performance, SSDs are magical to me because they are quiet. I am very willing to pay a premium just to be free of that drive seeking and grinding.
it's shouldn't hurt to use their numbers to guide your decision it is no guarantee though.
Living in the future is cool.
For personal drives in small quantities, you're basically operating on luck. You can get a bad HGST drive just like you can get a bad Seagate drive. I've had drives from HGST, WD, and Seagate and all of them have lasted far beyond the warranty expiration.
When I look into drives for small-scale personal use (NAS, desktop, etc), I check recent reviews about warranty and customer service, then buy based on price and features. I also buy drives based on the use case (NAS drives for NAS, desktop drives for desktop, etc).
You need to add "cost to repair" into your equation. At my old job we used to figure $300 to replace a drive (scheduling maintenance window with customer, taking backup, migrating/quiescing services), replacing drive, verifying services were happy, wiping or destroying old drive, managing RMA, qualifying/burning in new drive).
YMMV depending on policies and procedures. If I were Backblaze, I'd design the system to work equally well with 40 drives as with 45, and consider just leaving bad drives in until a handful of them need replacing in a chassis. And the replacement be basically automatic.
This version of the report is much closer than they've been in the past. Seems like Seagate might be making some improvements. IIRC, it was not uncommon for Seagates to be a 10x higher failure than HSGT in the past.
They designed their own replication scheme which allows quite a bit of disk to fail. Considering that, using RAID would be absurd so I'm pretty sure they are all configured as JBOD, thus allowing this kind of replacement.
While they do have multi-failure redundancy (17 data, 3 parity), it wouldn't be a good idea to push the limits regularly because that means you're risking the loss of actual customer data. And adding more parity to give you extra wiggle-room would result in both storage efficiency losses and performance losses. It probably wouldn't be advisable to expand the "stripe" size to cover an entire vault (to increase the independent disk failure tolerance without reducing the data-to-parity ratio) because you'd kill performance (the matrix multiplications necessary to do Reed-Solomon have polynomial complexity).
[1]: https://www.backblaze.com/blog/vault-cloud-storage-architect... [2]: https://www.backblaze.com/blog/reed-solomon/
> You need to add "cost to repair" into your equation. At my old job we used to figure $300 to replace a drive
This is TOTALLY true, and Backblaze does add in the cost to repair. At our scale, we have full time datacenter technicians that are onsite 7 days a week for 12 hours a day (overlapping shifts) so that we can replace drives and repair equipment within a certain number of hours.
Every single day our monitoring systems kick out a list of about 10 drives that need to be replaced, and the data center techs try to make sure there are zero failed drives when their shift ends. (We leave failed drives alone for up to 12 hours at night as long as it is just one failed drive in a redundant group of 20.)
But if you are operating with FEWER than our 110,656 drives then you can't have full time repair people standing around paying their salary. So you are really going to need to pay much more (per drive) than we do.
One of the absolutely amazing things about "cloud services" like Amazon S3 or Backblaze B2 is that (hopefully) we can actually operate it at a lower cost and higher durability than you can operate yourself. That may not be true at either end of the spectrum - if you have fewer than 10 drives it may be cheaper for you to operate it yourself, and if you have more than 1,200 drives (the deployment size of one of our vaults) you might want to consider cutting out the middle-man (Backblaze) and saving yourself money. But I make the bold claim we can save you money in that sweet middle ground.
One of the factors to consider is the cost of electricity in your area. Backblaze pays about 9 cents/kWh which is "pretty good". Up in Oregon you can get deals for 2 cents/kWh and beat our operating costs, but in Hawaii you are at 50 cents/kWh and unless you have solar panels you pretty much better host data on the mainland. :-) I think TONS of people (mistakenly) think that after they purchase a hard drive, operating it is "free". Electricity is one of Backblaze's major costs in providing the service. Our electrical bill is more than $1 million per year right now.
Some people think we (Backblaze) get "magical good deals" by purchasing drives in bulk, and it isn't as good as you might think. Sure, we get some bulk discounts, but think 5% or 10% better than your retail price for 1 unit. And you can probably get THE SAME DEAL as Backblaze if you are purchasing 1,000 drives in one purchase order.
Backblaze is a pretty good deal. I'm currently trying to figure out if I should take my video files that I'm editing for my personal movies and store them on Backblaze (transmission time being the issue there) vs. getting a 10TB drive to put them on or a ZFS array, vs. just deleting the source. The latter has a very competitive price. :-D
In some environments when a drive goes down, everyone goes into "red alert" mode and they have to rush to repair the drive or replace it. In our case we don't have to drop everything and run to it - so it helps keep the overall replacement costs down.