Google’s Reliability Team Sat Down for an AMA Right Before Gmail Exploded
techcrunch.com
techcrunch.com
I'm going to have to seriously think about the risks of being so heavily reliant on Google services.
I'll show them how it's done!
If you run a service yourself, you will have downtime, only you won't have Google's amazing SREs and wealth of experience to draw upon to resolve your outage. You will have to fix it yourself.
If you depend on some other service, you are subject to their uptime and disaster recovery capabilities. Nobody is better than Google at either of these things.
But to your point, it probably is a good idea anyways to have offline backups of your critical data, and to have out-of-band fallbacks for critical communication in the event your primary channels have an interruption.
crickets
"After a long and tricky debugging process, we found that a big MapReduce job was firing up every few hours and, as a part of its normal functioning, it was reading from /dev/random. When too many of the MapReduce workers landed on a machine, they were able read enough to deplete the randomness available on the entire machine. It was on these machines that our serving binaries were becoming unresponsive: they were blocking on reads of /dev/random!"
http://www.reddit.com/r/IAmA/comments/1w1y5m/we_are_the_goog...
Also, these guys are in engineering. They are very likely not even directly involved when there are outages. They build the systems and protocols to avoid and recover from outages, but don't actually perform the work themselves. It's developers vs. IT.
Correct, it's pretty much impossible for an outage to not be noticed and the GMail on-call being automatically paged.
SREs at GMail are engineers, yes, but they're very much directly involved with fixing outages - not so much at the 'try turning it off and then turning it on again' level, more the 'redirect all traffic away from this cluster into a different one, while we roll back the broken update'.
SRE is a combination of problem-solving when there are outages, and building tools to 1) automate away the manual jobs involved in massive-scale system administration so that outages are less likely to occur.
Doesn't sound like an organisation that would miss services going down even briefly.
Four SREs showed up. Two answered 4 questions each, another 8, last one 12.
Pretty poor to be honest.
Maybe Oracle?
That said, this was already posted in their original article about the Gmail downtime.
Now, Google could be better about this. It's the implication that they're somehow not good in the first place that's so off-putting.