Backblaze Durability Is Eleven 9s – And Why It Doesn’t Matter
backblaze.com
backblaze.com
Correlated failures are common in drives. That could be a power surge taking out a whole rack, a firmware bug in the drives making them stop working in the year 2038, an errant software engineer reformatting the wrong thing, etc.
When calculating your chance of failure, you have to include that, or your result is bogus.
Eg. Model A of drive has a failure rate of 1% per year, but when failed the symptom is failure of the drive to spin up from cold, however if already spinning it will keep working as normal.
3 years later, the datacenter goes down due to a grid power outage and a dispute with diesel suppliers so the generators go down. It's a controlled shutdown, so you believe no data is lost.
2 days later when grid power is back on, you boot everything back up, only to find out that 3% of drives have failed.
Not a problem. Our 17 out of 20 redundancy can recover up to 15% failure!
However, each customers data is split into files around 8MB, which are in turn split into the 20 redundancy chunks. Each customer stores say 1TB with you. That means each customer has ~100k files.
The chances that you only have 16 good drives for a file is about (0.97^16 * 0.03^4)2019*18 = 0.3%
Yet your customer has 100k files! The chance they can recover all their data is only (1-0.003)^100000... Which means every customer suffers data loss :-(
While I always find storage analysis interesting (I spent 5 years at NetApp where it was sort of a religion :-)) some of the assumptions that Brian was tossing out are not good ones to make. (like the lack of correlation, or that Drive Savers will exist as a company 10 years from now).
Still it does help you to understand the they take data availability seriously which is the underlying message.
> some of the assumptions that Brian was tossing out are not good ones to make.
We COMPLETELY welcome other analysis and listing other assumptions. Internally, we argued endlessly about why this or that wasn't totally accurate, and finally decided to publish the math WITH all of our assumptions exposed so you could be the judge. If Amazon wants to publish their assumptions for S3 for comparison, we're all ears.
> that Drive Savers will exist as a company 10 years from now
Absolutely true, this calculation is only good RIGHT NOW. For example, one of the things that came up internally was "well, when drives get more dense the rebuild time rises, so this calculation will no longer be accurate in two years". But at the same time, we have some additional tricks and optimizations to make which we have not done yet to cut the 6 day average drive rebuild time down to 3 days. Also, drives last us about 5 years, so your data will be migrated to totally new drives 5 years from now. Those drives will absolutely have a different drive failure rate (maybe higher, maybe lower) so the calculation will no longer have the same result 5 years from now.
As someone who likes to geek out on failure proof systems and perfectly secure systems, neither of which are attainable but can be asymptotically approached, I think you are seeing the "there is always one level deeper" kinds of discussions. Personally I think of them as endorsements because if the exceptions get too extreme (say 'what if an asteroid hits?') then you know you've got all the bases covered.
Its only a problem if the person analyzing the analysis finds something that you really did not even consider. Then it opens up an opportunity to look at the problem a whole new way.
[1] https://www.usenix.org/legacy/publications/library/proceedin...
The Cauchy matrices remove the need for a table lookup to do the Galois multiply, replacing it with pure XOR. Together with other optimizations, this gives nearly 3x-5x more coding throughput for the same (17,3) parameters, assuming you're still using your open-sourced JavaReedSolomon in production. I don't know if Reed Solomon coding throughput is a factor in your rebuild times?
It contains optimized Galois Field multiplication, resulting in Reed Solomon (both Vandermonde as well as Cauchy) at multiple GB/s on a modern x86 CPU.
Still, I doubt that the Reed Solomon coding speed is the limiting factor in their rebuild time. There is a mention of a 6-day duration, so even with a very slow Reed Solomon implementation ( ~ 100 MB/s) that should not be a bottleneck for a 10 TB drive rebuild (assuming a distributed rebuild approach, not a traditional RAID style rebuild).
The 2010 S3 calculation is obviously also wrong! I totally feel your pain in terms of wanting to have a directly comparable answer, but the reasons you give that "it doesn't matter" (and others, including correlated faults, software bugs, model risk, and security breaches) are actually reasons why the stated durability number is wrong. Honestly IMO a probability of 1-10^-11 is the wrong answer to pretty much any question; model risk is going to dominate that for any problem more complex than 1+1=2.
That said, although neither your system nor Amazon's should be expected to have anywhere near "eleven nines" of durability in reality, if as I understand it S3 is split across availability zones in a region and your product has all splits in a monolithic DC, I would expect S3 to come out ahead in a more careful analysis. (But note that S3 is not really a seamless multi region product, though there is an option to set up cross region replication.)
Though I'd welcome hard data on this very much.
This failure mode, at least, is already accounted for by sharding data across cabinets:
Each file is stored as 20 shards: 17 data shards and 3 parity shards. Because those shards are distributed across 20 storage pods in 20 cabinets, the Vault is resilient to the failure of a storage pod, or even a power loss to an entire cabinet.
https://www.backblaze.com/blog/vault-cloud-storage-architect...
However, they don't seem to offer multi-datacenter (or multi-region) redundancy so are still susceptible to a datacenter fire/failure.
In comparison, AWS S3 distributes data across 3 AZ's (datacenters), and you can further replicate across regions if you choose. Though you pay for that added redundancy in 3 - 4X higher cost.
From https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/Conce...:
> Each AWS Region has multiple, isolated locations known as Availability Zones.
And from https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/using-re...:
> Each region is completely independent. Each Availability Zone is isolated, but the Availability Zones in a region are connected through low-latency links.
This doesn’t list all the locations, but is a good map to get an idea:
https://www.google.com/maps/d/u/0/viewer?ll=50.9584270000000...
[0] https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/using-re...
AWS purposely doesn't publish that information, and while I can believe it's possible to crowdsource the data by doing a little sleuthing (or working for certain vendors), it's hard to trust the map without knowing the sources.
https://www.theatlantic.com/technology/archive/2016/01/amazo...
The AWS Cloud infrastructure is built around AWS Regions and Availability Zones. An AWS Region is a physical location in the world where we have multiple Availability Zones. Availability Zones consist of one or more discrete data centers, each with redundant power, networking, and connectivity, housed in separate facilities
https://docs.aws.amazon.com/aws-technical-content/latest/aws...
Am I reading that correctly?
More info here if you want to see how regions are structured at a high level of detail: https://youtu.be/AyOAjFNPAbA?t=15m21s
A chaos monkey that randomly powers down disks one at a time can prevent this.
You just made the same mistake you're criticizing. You assumed the 100K files were uniformly and independently spread. They're also likely clustered, and perhaps not even at the same data center. Given the variety of drives Backblaze uses, the drives are also not likely to all be the same model, so your failure method is also unlikely.
You also only looked at the case there are exactly 16 good drives. The proper failure estimate is 1 - (odds of 17 good + odds of 18 + odds of 19 + odds of 20). I'm not sure where you got the 20 x 19 x 18 part either. Did you mean 30 choose 16 or something like that? Using the proper 1-... method I get 0.00267, not 0.03.
Yes. Even if you take it into account, the vast majority of customers will see data loss, assuming random (but not even) shard distribution.
I rounded and approximated to avoid explaining too many probability rules... As you can see, it gives the same result at the end.
No, it doesn't. You picked 3% of all drives failing out of the blue. Your next estimate was an order of magnitude too high. Your last assumption of p^# drives is not reasonable.
The proof is in reality. Backblaze has run over a decade, with all sorts of hardware failures, server configs, running many drive models through their lifetime, across manufacturers, across technologies, across multiple datacenters, and had not seen the level of failures you claim they will.
So I suspect their method of estimating is more accurate than yours. So far it matches reality much better.
How many years between incidents are you talking, and when was the last time a manufacturer had a multigenerational bug? (and is quality improving, or decreasing? That is, are we more or less likely to see failures in the next 10 years than we did in the last 10?)
Disk failures could also be triggered by datacenter environmental factors shared among many drives like temperature or noise.
Your power outage causing 3% of drives to fail is just a subset of that.
Realistically, the chance of a tornado taking out the Swedish datacenter built inside a former nuclear bunker under 100ft of granite bedrock is so small that it probably doesn't affect the number of 9's that you can claim.
> Arctic stronghold of world’s seeds flooded after permafrost melts > It was designed as an impregnable deep-freeze to protect the world’s most precious seeds from any global disaster and ensure humanity’s food supply forever. But the Global Seed Vault, buried in a mountain deep inside the Arctic circle, has been breached after global warming produced extraordinary temperatures over the winter, sending meltwater gushing into the entrance tunnel.
There are always unforeseen and unforeseeable risks associated with any location. You can mitigate them but you can't claim X number of 9s for a single physical datacenter.
> There are always unforeseen and unforeseeable risks associated with any location. You can mitigate them but you can't claim X number of 9s for a single physical datacenter.
What is X, here? I'm pretty sure I can claim 99% for a single datacenter.
Nobody can promise 100%, but that doesn't mean that all those 9's are meaningless. They mean a lot for budgeting, and even more for insurance purposes -- which is exactly what we as a civilization have come up with as a way to amortize loss in the event of a local catastrophe. Your premiums are going to be much higher if you don't have enough 9's in a critical part of your money-making infrastructure.
No one here is saying that you don't need geographical redundancy. First we need to figure out how many 9's we can realistically expect in order to determine how much redundancy makes financial sense.
I mean, that's kind of what Backblaze is saying in the article, isn't it? They don't have geographical redundancy, yet there's not a single mention of that fact or the importance thereof in an entire article dedicated to teaching the unwashed masses about the limitations of mathematical theory in analyzing durability, even going so far as to say:
> somewhere around the 8th nine we start moving from practical to purely academic... it’s far more likely that...Earthquakes / floods / pests / or other events known as “Acts of God” destroy multiple data centers [emphasis my own]
Seems like a pretty serious omission given their claimed authority as "the bottom line for data durability" and being "like all the other serious cloud providers" who do have geo redundancy, don't ya think?
If you need to be safe about your data, you NEED several cloud providers in different places, with different softwares and different countries.
Especially for data it is pretty easy to just back it in 2 really different places at different providers. Relying on the geo redundancy of ONE provider and having to pay for it seems a bit useless for me.
Personally, I don't care whether a single provider has multiple datacenters or not, because I prefer to have redundancy across providers. But that's not the kind of recommendation that we're likely to see on the blog of one of those providers.
Which means every customer suffers data loss
They suffer partial backup loss.The customer only suffers data loss if they lost their 'master' copy of the data as well during the outage. Iff they don't have a secondary backup solution.
If it's the latter, "data loss" seems an appropriate characterization.
This is why, when I was building DIY arrays for startups (around the same time Backblaze published their first pod design [1]), I went through the extra effort of sourcing disks from as many different vendors as possible.
Although it was somewhat more time consuming and limited how good a price I could get and how fast the delivery could be, it meant that, for any given disk drive size, it meant I could build an array as large as 12 where no 2 drives were identical in model and manufacturing batch [2].
Of course, it's still a vanishingly rare risk, and "nobody" cares about hardware any more. It does help to remember, at least once in a while, that, on some level, cloud computing really is "someone else's servers" and to hope that someone else still maintains this expertise.
[1] though I used SuperMicro SAS expander backplane chasses for performance reasons
[2] and firmware from the factory, although this is somewhat irrelevant, as one can explicitly load specific firmware versions, and, IIRC, the advantages of consistent firmware across drives, behind a hardware RAID card, outweighed the disadvantages
This is very good advice!
If you already built your array, consider advice: "replace a bad disk with a different brand, whenever possible".
Over time, you naturally migrate away from the bad vendors/models/batches. After following this practice, it seems ridiculous to me now to keep replacing the same bad disks with the same vendor+model.
Some of this can also be achieved ahead of time if one has multiple arrays with hot spares, by shuffling hot spares around, assuming there's some model diversity between the arrays but not within them.
I doubt I'll ever again have the luxury of being able to perform this kind of engineering, however. Even a minor increase in cost or cognitive/procedure complexity or a decrease in convenience just serves to encourage a "let's move everything to the cloud" reaction, so I keep my mouth shut.
I assume you're talking about already-written sectors becoming unreadable or a similar failure. Unfortunately, I don't think you can. This is what I believe the "patrol read" feature of RAID cards is meant to address.
Fortunately, however, I don't believe there's evidence that if the data is readable, it would ever be different from what had been written, so comparison isn't needed. The main exception to this is the case of firmware bugs that return sectors full of all-zeros.
> Does that impact durability of the drives?
I haven't read the studies (from Google, mostly, IIRC) in a while, and I'm not sure if they've released anything lately for more modern drives [1]. However, I believe you'll find an occasional "patrol read" won't noticeably reduce drive life/durability.
[1] Especially for something like SMR, whose tradeoffs would seem particularly attractive for something like this archival-like use case.
Comparison is needed to address misdirected writes and bit rot in the very least, see "An Analysis of Data Corruption in the Storage Stack" [1]. You can't count on your drive firmware or RAID firmware to get this right. You need bigger end-to-end checksums, and you need to scrub.
Still:
> On average, each disk developed 0.26 checksum mismatches.
> The maximum number of mismatches observed for any single drive is 33,000.
Considering the latter can represent 132K on a modern, 4K-sectored drive, that's a remarkable amount of data loss, enough to warrant a checksumming higher up (such as in the filesystem).. in theory [1].
However, the fact that this was NetApp using their custom hardware as the testbed makes me wonder if the data are skewed, and if the numbers would be nearly this bad from a more "commodity" setup, such as at Google. The paper alludes to this when referring to the extra hardware for the "nearline" disks, and I'm always suspicious of huge discrepancies in statistics between "enterprise" and other disks, even more so when there's a drastic difference in comparison methodology.
It would be interesting to see if there are any numbers for more modern drives, especially as the distinction between "enterprise" and "consumer" drives is disappearing, if only because demand for the latter is disappearing.
[1] In practice, an individual isn't aware of the 16.5K/132K loss risk, which is vanishingly small compared to other risks, anyway, and businesses don't tend to care and have survived OK anyway.
Higher-level checksum failures are, however, a situation, where I would appreciate an integration between filesystem and RAID, as I'd want a checksum error to mark a drive as bad, just like any other read error.
Do you happen to know if ZFS does that?
This is super-disturbing and a dealbreaker, if it's still true:
> The scrub operation will consume any and all I/O resources on the system (there are supposed to be throttles in place, but I’ve yet to see them work effectively), so you definitely want to run it when you’re system isn’t busy servicing your customers.
I browsed a little of the oracle.com ZFS documentation but couldn't find much in the way of what triggers it to decide that a device is "faulted" other than being totally unreachable.
But if you have faulty power or bad hardware chipping away at your equipment, your depth of resiliency is degraded until the issue is corrected.
Your math is completely wrong. In reality 92% of customers suffer no data loss.
The chance of a file having 4 failed drives (16 good drives) is: .03^4 = 0.00008100%
The chance of a file being irrecoverable is the chance of having 4 or more failed drives: .03^4 + .03^5 + .03^6 + ... = 0.00008351%
The chance of a file being recoverable is: 1 - .03^4 - .03^5 - .03^6 - ... = 99.99992151%
The chance of a customer's 100k files all being recoverable is: (1 - .03^4 - .03^5 - .03^6 - ...)^100000 = 92.0%
Therefore only 8% of customers encounter one or more 8MB file that is irrecoverable.
The OP still has a slight error though. C(20,4) = 4845 possible combinations of failing drives (4 out of 20.)
Therefore the chance a file is irrecoverable is: .97^16 × .03^4 × 4845 = 0.241% (not 0.3%)
But the OP's conclusion is largely correct: every customer will have some irrecoverable files.
Edit: actually the math is still wrong. The chance any 4 out of 20 drive is failing is: .03^4 × C(20,4) = .03^4 × 4845 = 0.392% — There is no need to multiply by .97^16 as the status of the other 16 drives is irrelevant.
• 00000000000000000000 = all 20 drives are working
• 00000000000000000001 = D19-D1 working, D0 failing
• 00000000000000000010 = D19-D2 working, D1 failing, D0 working
• 00000000000000000011 = D19-D2 working, D1-D0 failing
• etc
The probability of each of these scenarios is:
• 00000000000000000000: .97^20
• 00000000000000000001: .97^19 × .03
• 00000000000000000010: .97^19 × .03
• 00000000000000000011: .97^18 × .03^2
• etc
There are C(20,4) = 4845 scenarios with exactly four failing drives (four "1" bits.) The probability of each scenario is .97^16 × .03^4. Therefore the probability of 4 failing drives (any drive) is the sum of the probability of each scenario: .97^16 × .03^4 × C(20,4) like I said 3 comments above.
However the probability of a file being irrecoverable is P(4 failing drives) + P(5 failing drives) + ... + P(20 failing drives):
.97^16 × .03^4 × C(20,4)
+ .97^15 × .03^5 × C(20,5)
+ ...
+ .97^0 × .03^20 × C(20,20)
= 0.267%https://www.usenix.org/legacy/event/fast08/tech/full_papers/...
http://www.cs.toronto.edu/~bianca/papers/fast07.pdf
Take note of section 5 in both papers for their statistical models and how they match with real-world data.
I run into people all the time that seem to have the same problem, to the point that it makes me wary of any software developer putting forth numbers that seem fantastical.
My gut reaction is that if Backblaze wants to keep reporting their disaster preparedness numbers that they need the assistance of an actuary to calculate them.
Because at these probability levels, it’s far more likely that:
- An armed conflict takes out data center(s). - Earthquakes / floods / pests / or other events known as “Acts of God” destroy multiple data centers. - There’s a prolonged billing problem and your account data is deleted.
The point that once you get to a certain point of durability (at least as far as hardware/software is concerned) you're chasing diminishing returns for any improvement. But the risks that are still there (and have been big issues for people lately) are billing issues. I think it's an important point that the operational procedures (even in non-technical areas like billing and support) are critical factors in data "durability"
This seems like a form of the same argument, and I wonder where else this arises and how people describe it.
Specifically: every X-nines durability design will be compromised by some failure mode you didn’t think of.
For example, in the hash collision case the argument says that it's not worth worrying about the (known) probability of a software error due to an unexpected hash collision because it's dominated by the (known) probability of a comparable error due to cosmic radiation. (The former probability can be calculated using the birthday paradox formula, and the latter has been characterized experimentally in different kinds of semiconductor chips.)
This kind of argument doesn't rely on the idea that there are other risks that we can't identify or quantify. It's about comparing two failure modes that we did think of, in order to argue that one of them is acceptable or at least not worth further attempts to mitigate.
The point about billing is better, though.
My other concern is that a software bug, operator error, or malicious operator deletes your data.
But if you're choosing two data points, my question is... why two? If you are choosing whether or not to reply based on whether or not the second data point fits with the first, then you're introducing selection bias. The chance that the second data point disagrees with the first by at least as much as the 200 My interval disagrees with the 1/65 My rate is equal to 1-(exp(-65/200)-exp(200/65)) = 0.32, which is not especially high.
N, of course, is always 1.
[t/3, 3t] with 50% confidence
[t/4, 4t] with 60% confidence
[t/39, 39t] with 95% confidenceI don't think this is a useful definition of your yearly durability. If your data center is down for maintenance during a period in which it is guaranteed that nobody wants to access it, that doesn't reduce your availability at all -- if your only failure is an asteroid that kills all of your customers, it would be more accurate to say you have 100.000000% availability than 99.999999%.
They usually give you a reasonable amount of free storage, and it's unlikely all accounts would be terminated or locked at the same time.
And of course, you should always have your local backups as well.
Replace with Amazon and/or crashplan as appropriate..
[1] https://www.computerweekly.com/opinion/Nirvanix-failure-a-bl...
1) Sideloading. I was unable to benchmark or even to get this to work. Images requested to be loaded from media-src.nowpublic.com to node4 @ Nirvanix never showed up. I showed my code to [nirvanixcontact] who said that the code looked OK but someone did a load testing on node4 without informing Nirvanix and that steps are made so that such a situation won't occur again. He told me that the requests are not lsot but they still have not landed. Again, my code can be at fault and I would be happy to run a sideload example code.
2) Upload speeds. I have uploaded from d2 to Nirvanix 100 images each between 180-8584 kbytes totalling almost exactly 100 MB (101 844 931 bytes). The upload was a single HTTP request. The uploads took 18-19 minutes (I repeated the experiment). To give us a comparison I changed the URL in the script to a one line PHP script on another server (at hostignition.com) which just echo'd the number of uploaded files. This took 16.25 seconds and echo'd 100 so seemingly the files landed.
2a) I tried to get another node via the LoginProxy method which we would need for uploads anyways. While LoginProxy itself did work, GetStorageNodeExtended https://services.nirvanix.com/ws/IMFS/GetStorageNodeExtended... always fails with ResponseCode 80006, ErrorMessage: Session not found for ip = 67.15.102.70, token = e7b00d25-fc35-431c-9437-9a4302767f46. Seemingly, does not pick up the consumerIP.
3) The image conversion itself is blazing fast though. These images took a total of 101.58 seconds to convert and this includes 300 HTTP requests (200 sent to d2 from Nirvanix, 100 to Nirvanix).
You might get killed in the course of the initial police raid though..
(kids below age of consent sexting each other is another, related, problem)
You have horror stories like https://www.reddit.com/r/tifu/comments/8kvias/tifu_by_gettin...
> Eventually someone realized that their non-work accounts were banned as well. It wasn't until yesterday that someone made the connection. Anyone who had their accounts as a recovery option were also caught in the ban wave.
(this isn't quite true in my case, but Google did go to some lebghts to merge yt accounts into Google accounts recently).
I don't recall hearing about anyone recovering access to their data in that case?
This makes every additional user an increase in risk, as even without a warrant it seems if USA TLAs consider someone a valid target then those servers are going down (or getting taken over by people with unknown service standards in order to run a sting, or ...).
tl;dr you have to worry about accusations (of law breaking or copyright infringement) against others too as some jurisdictions have a strong overreach in such cases.
I've written in more detail before[0], but just to share the gotchas in case anyone here is thinking of switching to Backblaze:
1. They backup almost no file metadata.
2. The client is very slow (days or more) to add new files and there's no transparency (it claims everything is backed up when it's not).
3. There are still bugs in the client that can put your backup into an invalid state where it gets deleted.
4. Support is terrible, and won't be any help when you run into these bugs.
I've been using rclone (got the recommendation here) which has been reliable.
Also, does anyone know if Backblaze has any plans to offer u2f? I've switched dns and email providers to get u2f.
And I don't know if it supports U2F, but it does support TOTP.
[0] https://mjtsai.com/blog/2014/05/22/what-backblaze-doesnt-bac...
Their pricing is amazing, but saving money on a back-up solution that doesn't seem as good as the other cloud storage providers is a dangerous game.
That's why you need another solution if you are serious about your data, maybe a set of external hard drives (local backup). This way, you have redundancy and little correlation in failure, which greatly improves your general durability. That local storage may be paid with the money you save by getting an "inferior" cloud backup provider.
> 30MBit/s sustained for 3 days isn't unreasonable for a business
We (Backblaze) are seeing more and more consumer internet connections in the USA with 20 Mbit/sec upstreams, I thought they were available most everywhere if you were willing to upgrade your internet package "just a bit". 30 Mbits is a little unusual for "consumer", but not unheard of. Of course, there is a "selection bias" when you look at online backup users. :-)
From Comcast: The Terabyte Internet Data Usage Plan is a new data usage plan for XFINITY Internet service that provides you with a terabyte (1 TB or 1024 GB) of Internet data usage each month as part of your monthly service. If you choose to use more than 1 TB in a month, we will automatically add blocks of 50 GB to your account for an additional fee of $10 each. Your charges, however, will not exceed $200 each month, no matter how much you use. And, we're offering you two courtesy months, so you will not be billed the first two times you exceed a terabyte.
Also, All customers in locations with an Internet Data Usage Plan receive a terabyte per month, regardless of their Internet tier of service. and The data usage plan does not currently apply to XFINITY Internet customers on our Gigabit Pro tier of service. The plan also does not apply to Business Internet customers, customers on Bulk Internet agreements, and customers with Prepaid Internet.
Backblaze is actually faster than Apple Time Machine is on my LAN which slightly bothers me. It also has lower CPU usage.
I originally chose Backblaze after benchmarking the other offers available at the time (Carbonite, Crashplan etc) and Backblaze was by far the fastest.
I found the discussion about why it doesn't matter when you start talking about 11 nines of reliability to be hilariously true.
At the end of the day we're still flawed humans living in a hostile universe, and no matter how foolproof we make the technology, there are some weaknesses that just can't be eliminated.
> well written and refreshingly transparent
Thank you! In the interests of full transparency, the blog post was a collaborative affair and was proof read and edited for clarity by several people at Backblaze.
> discussion about why it doesn't matter
One of the philosophies Backblaze uses is to build a reliable component out of several inexpensive and unrelated components. So combine 20 cheap drives into a ultra reliable vault. We have two or three inexpensive network connections into each datacenter instead of buying one REALLY expensive connection for 8x the price. Etc.
Personally, I recommend customers do the same. Instead of storing two copies of your data in two regions in Amazon for 2x the price, store your data in one region of Amazon and put one copy in Backblaze B2 for 1.25x the price. We believe this will result in higher availability and higher durability that two copies in Amazon because Amazon S3 and Backblaze B2 don't share datacenters in common (that we know of), don't share network links, don't share the software stack, etc. For bonus points, use different credit cards to pay for each, and have different IT people's credentials (alert email address) on each. That way if one IT person leaves your company and you don't get an alert that the credit card has expired, hopefully your other copy will be Ok.
However, there are risks that are not necessarily independent such as the US Government ordering these two services to delete your data, or, as the article mentions, an armed conflict destroying data centers.
Putting your data in either Backblaze B2 or Amazon S3 suffer from other failure modes outside of the durability of the raw system. For example, let's say your IT person is poking around in their Amazon S3 account and accidentally clicks the "delete" button and all your data is gone? Or what if your credit card has a transaction declined, and your IT guy has left your company or the emails from Amazon are being put in the "Spam" folder of your email program. Or maybe a malicious Amazon employee writes a program to delete all the data in Amazon S3 from all customers? What if one of your employees is really disgruntled and logs into your Amazon S3 account and just to spite you deletes all your data?
In every one of these situations, if you have a copy in Backblaze B2 and also another copy in Amazon S3, you can recover your data from the other vendor.
I recommend using a separate credit card to pay for your Amazon S3 account and your Backblaze B2 account. They should expire a year apart. And don't give the logins to both systems to one disgruntled employee in your organization. Only give that disgruntled employee access to one or the other.
Make sense?
Recovery workload should be spread across the whole cluster, so that the recovered data gets distributed evenly. In that case, assuming 10,000 drives, to recover one dead 12TB drive and a recovery rate of even 10 MB/secs per machine, recovery of one drive should be done in under a second. Maybe 10 seconds with some sluggish tail machines.
Why do you need it done in under a second? While the data is down one replica, it is at dramatically higher risk. Also, drive failures can be dramatically accelerated, for example in the case of a bad software release erasing data - you need to be able to move data faster than bad software gets released. And releasing software at a rate of one machine per second still means a release takes 3 hours!
In most systems we assume that "primary traffic" (read/write stuff) is prioritized over "rebuild traffic" which is recovering lost shards. So when you specify these things it is best to specify "how long to rebuild a shard while the array is providing storage services at its maximum specified rate." This assures the customer that if they have a 24/7/365 non-stop traffic pattern their data will still stay protected in the face of drive failures.
> During the 6 days the data would be available it just might have to be reconstructed on the fly by the error correcting rather than read directly.
Correct. More specifically, the FIRST time the data is accessed in any 24 hour period it must ALWAYS be reconstructed from the Reed-Solomon encoded parts on 17 other drives on 17 other machines. Any 17 is fine, so it's totally fine if 1 or 2 drives are not available. Once reconstructed it is stored in a set of front end cache computers that have fast SSDs for this purpose.
The second time the same file is accessed in a 24 hour period, it will be fetched out of the SSD cache layer so it won't even hit the spinning drives and won't care if all 20 drives are offline.
> "primary traffic" (read/write stuff) is prioritized over "rebuild traffic"
Yes. Backblaze balances between the two if only one drive has failed, but as a tome (20 drive group spread across 20 computers) becomes more badly degraded Backblaze begins favoring the rebuild. When two drives have failed out of 20, Backblaze stops allowing any writes to that tome because more writes will tend to fail yet another drive. Fewer writes offloads the tome. But we still allow reads. At Backblaze, we have never been 3 drives degraded out of 20 (knock on wood), but if this ever occurs the 20 drive tome is now running without parity -> so in that case we even stop allowing reads AT ALL until we are returned to at least 1 drive of fully redundant parity.
I want to know where you can find a drive that can write 12TB/sec of data!
(In other words, you clearly missed half the problem. To add a new replacement drive, you have to be able to write to it the data from an original drive. Also RS code calculation is fast these days, but it ain’t that fast)
> spread across 10,000 drives, which is unrealistic
I claim it is also undesirable. Backblaze specifically made the conscious decision that the parts of any one single "large file" (these can be up to 10 TBytes each) are all stored within the same "vault". A vault is 20 computers in 20 separate racks. This allows a single vault to check the consistency and integrity of a large file periodically without communicating to other vaults in the datacenter.
The vaults have been a really good unit of scaling for Backblaze. If the vaults can maintain their performance, then we know we can just stamp out more vaults because there is almost no communication between vaults.
But spreading the data over 10k drives isn't unrealistic, it's a different architecture. Pick a different 20 drives for each file.
Working it through: Assume 200 machines with 50 drives each. Each machine has to read 1TB, transmit it over the network, do a parity calculation, and write out ~50GB. With dual 10gbps ports the bottleneck is the network, and if we dedicate one on each machine to the rebuild we get a 15 minute clock.
Not that having such a monolithic architecture is worth the complication and extra bugs.
It will also be recovered by reading the recovery data, which should also be approximately evenly spread across the 10,000 drives.
Since others have already demonstrated why the remainder of your comment is overly simplistic, I'll tackle this bit.
Generally, software releases are not rolled out at a constant rate to all machines. A typical thing to do is to release it to staging, then to a "canary" subset of machines (e.g. to 1% or 5% of the machines).. Once all seems well there (e.g. metrics are clean and the canaries have handled X writes, reads, and simulated drive failures), it can be rolled out to a larger subset, and eventually to all machines.
In that way, the release can take whatever total amount of time is desired while still catching any such bugs fairly reliably.
Ideally, at backlblaze they could ensure that their canary instances are "data-redundancy aware" such that even if the 5% they roll to for the canary test all explode, data is still safe.
Regardless, any talk of "recovering data faster than software releases" is completely silly and totally misses the reality of how releases are done, how recovery is done, and what sort of bugs might happen. The math based on faulty assumptions about rate is also pointless.
Can this be reformulated: you store 10 trln objects (e.g. 100TB of 10 byte records), you lose 1 record each year.
Also curious what are the stats from other providers.
Roughly, yes.
> Also curious what are the stats from other providers.
As to other providers, most are 6+ 9s that I’ve looked at, with many in the 8-9 range. Anything over 8 is (as they admitted) essentially marketing porn and not a useful metric (for reasons they mentioned as well as ones said by other comments here).
I think a much bigger scandal is that all major laptop Operating System vendors (Microsoft and Apple) absolutely know when your laptop drive loses files or even is starting to go bad in some cases, and they NEVER tell the customer. I think an excellent product offering would be a 3rd party piece of software and cloud service which was a "verification service". It wouldn't store your files offsite, it would store the name, size, and SHA1 offsite and periodically check that no bits have been flipped on your local drive unless you intended it. For example, a week after I take a photo, I absolutely never want the photo to change. Ever. Same with music I (legally) download.
locally redundant storage: 99.999999999 % (11 9's)
zone redundant storage: 99.9999999999 % (12 9's)
geographically redundant storage: 99.99999999999999 % (16 9's)
Chunks of your data are going to be stored together, so it's a very small chance of losing a big block of 10 byte files. There's no failure mode that loses just one, and does so often.
It probably depends on their infrastructure, e.g. if storage is something like cassandra, records would be evenly distributed by key hash.
I agree that numbers likely are not like that though, just wanted to demonstrate that such calculation approach can bring unexpected conclusions.
Oh gosh thanks Backblaze, I'll just dig through several TB of stuff....
You need to do the same calculation for your meta data, which is probably not erasure coded. If you lose this, you don't lose your data, but you no longer know where you put it.
So you probably add your meta data to your data as well in some kind of recoverable format. That's fine, it means that you can harvest the meta data again.
But how long does this take ?
(side note: Werner is a great person)
Unfortunately, there is a difference, a huge difference, between a system "designed" for 11 9s of durability, and a system "offering" 11 9s of durability.
I wish Backblaze, or Amazon, or anybody else, would clarify durability using very honest terms.
An example?
"This system offers X 9s of durability over a period of one year, on average. This is a technical paper that describes how we tested that durability", followed by measurements and test specifics.
Any other claim has much less value to me.
I wonder how many times they've had 2 or 3 drives fail before they've rebuilt and if that matches their predictions.
Well written, but there are other significant risks like losing access credentials (e.g. a password stored only on one device that is destroyed in the same accident in which its only user, who remembered the password by heart, dies) or being hacked by someone who gains access to cloud storage and intentionally erases or corrupts data.
Specialization is good, but if Backblaze is strictly in the business of storing data on hard disks, who's going to help with designing and maintaining the reliable complete system on top of their service that users actually need?
In addition to calling (and possibly getting blocked/ignored), does your customer service staff send text messages? I suspect that a big percentage of the phone numbers you have are for cell phones these days, and I see a lot less SMS spam than I do telemarketing. SMS would also allow you to get a bit of info visible to recipients (e.g. "Backblaze CC Expired") with more detail once a message is opened.
https://www.nytimes.com/interactive/2018/05/24/us/disasters-...
As someone else pointed out, it’s overly simplistic. They’re a great low-cost alternative to S3, sure. But keep a backup on another continent if you need your data 100 years from now.
They don't have any way to detect corruption in the data or if they have, the backup clients are oblivious to it.
I lost about a 150GB of family photos and videos.
Minor nitpick, this ignores the possibility of more than 4 failures, although this error only affects the fourth digit after the nines. Much more egregious is the following:
>there are 56 “156 hour intervals” in a given year
This is too simplistic, there are in fact infinitely many 156-hour intervals in a year, some of them just happen to overlap. This overlap can't simply be ignored because even if none of their 56 disjoint intervals contain 4 events this does not rule out the possibility of there being 4 events in some 156 hour interval they didn't take into account. In fact failing to take into account even one of the infinitely many intervals creates a blind spot (consider what happens if the drives happen to fail precisely at the start and end of a particular interval). You can still get a lower bound by e.g. ensuring none of the 56 intervals contain more than 1 failure, or by adding more intervals and ensuring none of them have more than 2 failures etc.
Their binomial calculation contains the same mistake.
A quick improved lower bound can be obtained by calculating the probability that any failure is followed by (at least) 3 other failures within 156 hours. For one failure this probability is given by the Poisson distribution and is
Pc = 1 -\sum_{k<3} e^-λ λ^k / k! = 5.18413e-10.
Now we get into some trouble because the failures and the probability of a 'catastrophic' failure are dependent, however the probability that any particular failure turns catastrophic is constant, so the expected number of catastrophic failures can't be greater than the expected number of failures times that constant, this gives a lower bound of Pc (365·24·λ) = 6.63154e-9
this is a lower bound, but that's still three fewer nines left than their claim.Anyway let's just hope their data centres are more reliable than their statistics.
Edit: This last calculation can be justified by noting that the probability that 1 critical failure starts in a particular time interval is Pc times the probability of 1 failure in that interval plus some constant times the probability of more than one interval. Similarly the probability of more than one critical failure is at most the probability of more than one failure.
Now the probability of more than one failure in a time interval is dominated by the length of the interval, therefore if you calculate the density those parts fall away and you're left with a density of Pc λ critical failures per hour.
This seems to be an exact expression for the expected number of critical failures, and not just a lower bound. Although it is still a lower bound for the probability of a critical failure, albeit a fairly tight one.
Hum, I would be more reassured by past statistics than a probability evaluation. Did they happen to have loss data since their creation?