Google loses data as lightning strikes
bbc.com
bbc.com
I think my favorite part of that is "on average", as if you will be making repeated ten-million-year trials of this effectively brand new technology.
The point is that once you get into several nines of reliability, really rare events that are impossible to model start to dominate your risk budget.
So Google just exhausted their 11 9's for centuries to come.
To accurately compare them you'd need to look at AWS EBS: "Amazon EBS volumes are designed for an annual failure rate (AFR) of between 0.1% - 0.2%, where failure refers to a complete or partial loss of the volume, depending on the size and performance of the volume."
" This outage is wholly Google's responsibility. However, we would like to take this opportunity to highlight an important reminder for our customers: GCE instances and Persistent Disks within a zone exist in a single Google datacenter and are therefore unavoidably vulnerable to datacenter-scale disasters. Customers who need maximum availability should be prepared to switch their operations to another GCE zone. For maximum durability we recommend GCE snapshots and Google Cloud Storage as resilient, geographically replicated repositories for your data. "
In a cloud based infra, one still needs to know the difference. And for the ultra-low numbers, they are from mathematics I think. (# replicas, possibly cross DC/geo)
Your main point (that making high reliability systems more conventionally reliable is building your high wall yet higher still) is definitely valid. But the lossage rate is actually a meaningful number given the extremely large number of objects stored in S3.
Our data integrity is such that it will survive universal cataclysm and the reduction of all data to a single point of infinite energy.[1]
[1]guarantee applies to a single bit of data[2] [2]bit value must be 1
The Long Now foundation (http://longnow.org) has some other interesting projects besides the clock.
If you like this kind of thing, I'd suggest reading Neal Stephenson's Anathem. It might be hard to get into it, it might be his most impenetrable book yet but it really pays off later.
http://journals.aps.org/prd/abstract/10.1103/PhysRevD.32.248... (S. W. Hawking, 1985 <- Just so you know it isn't a crackpot writing this. It isn't a widely accepted view though)
But those nines are only true if Amazon's software is bug-free. I've lost data on Glacier on my home account. They personally called me to apologize.
Heck once in every 10,000,000 years, the odds of winning the USA lottery are around one in 175,000,000 and yet people partake and somebody wins. So it is always worth looking at odds from another perspective and accept it is either impossible, probably or possible and as we know the return from investment in durability is logarithmic in cost going up and return getting smaller.
Still one of the biggest issues in any datacentre is not just the power, but also the quality of the power and even the slightest noise could and does increase the odds of some electronics going wrong. UPS's (though been many years) will put out a square wave in modulation and the ideal is a perfect sine-wave. It is truly educational to see the quality of many power sources on an oscilloscope. Also equipment on the same circuit can also induce noise, so can vary rack to rack.
(Forgetting about correlation was a big part of the MBS and LTCM financial failures)
Assumptions that variables are independent are often very mistaken.
I wouldn't trust any of these figures unless they have ongoing efforts to test them empirically. E.g. create distributed databases of 100 trillion objects, mess with them in various ways, and perform correctness checks on them.
Also, the typical civilization-ending comet or asteroid is going to destroy 0 or 1 geographically-diverse data centers. Remember that this is a retention promise, not an uptime promise.
If you're proposed reliability is data loss of one in 100 billion per year, that assumes the risk of e.g. global thermonuclear war is no more than that amount, or you're ignoring it for the purposes of the calculation. Which is why I wonder about their assumptions.
As for a comet destroying at most one data center, if it really ends civilization then the data will probably be lost before too much longer even if the facility physically survives.
When you are storing many billions of files consuming many millions of disks, you become vulnerable to black swan events, simply because you are rolling the dice so many times.
Important detail left out.
If there are 10,000,000 projects with 10,000 objects, then a system-wide durability of 99.999999999% would expect to drop 1 object per year.
[0] http://science.ksc.nasa.gov/shuttle/missions/51-l/docs/roger...
Ever thought about how we can talk about the chance of rain tomorrow? Or the risc that a comet may strike the earth within the next million years?
The surge suppression gear we put in (lead ins at power feeds, RF feed, etc) is mostly to prevent a fire and to ensure the extra energy goes largely to ground.. but it won't prevent dead gear.
EDIT:
All right, I'll rephrase. According to Google's infobox from nat'l geographic, lightning generates up to 1 billion volts.
-> Are surge protectors at even the highest-end data centers simply not rated to a billion volts of surge protection?
That's an OK use of the language in this context and has many parallels in other fields. E.g. using a condom as protection against pregnancy and disease. Many people have learned the hard way that it's not 100%.
I use an isobar myself to power reduce noise, but I'm under no illusion that it will protect my shit from a direct strike.
There are a lot of cloud-to-cloud backup services out there, but to me that seems like the blind leading the blind, especially with regards to malicious data destruction. For instance, I've recently been experimenting with Cloudally to automatically back up Google Drive, which seems like a good solution at first- until you think about the fact that Cloudally uses Google accounts for authentication (and doesn't use 2FA for native authentication). In other words, an attacker with access to my primary data (Google Drive) would also have access to my backups. Worse than that, Cloudally actually increases the attack surface, since its lack of 2FA presumably makes it easier to crack than my Google account.
Similarly, I'm guessing a lot of cloud backup services share data centers with the services they are backing up.
This is how, for example, you can know whether a wildfire was started by lightning - once a point of origin is determined, simply check the data for strikes.
https://en.wikipedia.org/wiki/Lightning_detection http://www.lightningmaps.org/
Assuming 1 petabyte of total storage at the datacenter, that equates to about 100mb. I wonder how much storage they have there.
XKCD guessed 15EB in 2013
I thought battery is supposed to cover writing the entire write buffer cache to disk in case of power loss. Sounds like they had some badly designed gear which did no account for partial battery charge which should downsize the cache to battery's capacity.
"In a very small fraction of cases (less than 0.000001% of PD space in europe-west1-b), there was permanent data loss."
https://status.cloud.google.com/incident/compute/15056#57195...
Did anyone else cringe?
I use google drive a lot, I don't track what's in my drive, should I be worried?
Achieving RPO=0 generally requires synchronous replication to a different datacenter which adds significant latency.
OTOH, Google does datacenter redundancy across different locales, which make synchronous replication perform much worse, like you noted.
"Wouldn't successive lightning strikes be less likely due to reduced local electric potential?"
It's not readily apparent to me which of the effects will be larger.
What if I edit the same pixel, in the same image, in the same nanosecond on two different devices, and then sync?
There are situations where conflicts are necessary. The focus should be on making conflict resolution easily accessible, not on trying to be smart and overriding files at random
Surely this is why Google docs can manage collaborative working more easily -- because they can control of the whole stack.
Better than the OP would be file syncing of arbitrary files between arbitrary systems is hard.
(Disclaimer: I work there but have absolutely no knowledge about how Drive stores data)
I'm sure the data wasn't truly lost. If I'd called them up and they'd made it their priority to find my old files, they could have done so, having so many redundant backups. But of course no one at Google is taking calls like that or acting on individual requests. The data was effectively lost, not technically lost. But I'm sure it's uncommon.
Of course once you get beyond the headline, I think most people are much worse with protecting themselves from rare outages than Google.
Normally, Google redundantly distributes out data to at least 3 different geographically distinct locations. Check out the 'BigTable' white paper [0] for more info.
For 99% of cases (and pretty well all user cases), this would not cause data lose. The key here is that the data was generated on the servers and did not have a chance to duplicate before the event.
link here: http://static.googleusercontent.com/external_content/untrust...
actually in colossus one can tune RS coding parameters per file, to get a tradeoff between performance/durablity.
RS coding uses less copies, but same level of safety (tradeoff is the recovery computation time.)
EDIT: In this video https://vimeo.com/100153741, around the 23 minute mark
Unless I am specifically paying for backup service, I wouldn't expect/want them to do that for me. And even if I was, I'd still have offsite backups if the system was important.
If you aren't being responsible with your backups in a noisy environment like Google Cloud/AWS, understand that you are vulnerable to freak accidents like this. Google/AWS's job in all of this is to try to reduce the frequency of issues and to minimize the impact.
Google Compute Engine offers customers the option to make snapshots for backup, or use a true "cloud" storage engine. If anyone lost data here it was customers explicitly not doing backups and only using a single zone. I don't know why anyone would expect different. GCE easily allows you to network machines in multiple data centers, but close geographically. So you'd only need to handle region-wide disasters.