Google Cloud Networking Incident Postmortem
status.cloud.google.com
status.cloud.google.com
Answer, and the root cause summarized:
Maintenance started in a physical location, and then "... the automation software created a list of jobs to deschedule in that physical location, which included the logical clusters running network control jobs. Those logical clusters also included network control jobs in other physical locations."
So the automation equivalent of a human driven command that says "deschedule these core jobs in another region".
Maybe someone needs to write a paper on Fault tolerance in the presence of Byzantine Automations (Joke. There was a satirical note on this subject posted here yesterday.)
At some point people realized multi-server systems within one AZ are prone to failure. They then started deploying their system redundantly to multiple AZs within the same region to increase availability. This helped, but created consistency issues. To fix this people started building multi-AZ software systems, creating dependencies across AZs that weakened overall availability.
At some point people realized multi-AZ systems within one region are prone to failure. They then started deploying their system redundantly to multiple regions of the same cloud platform to increase availability. This helped, but created consistency issues. To fix this people started building multi-region software systems, creating dependencies across regions that weakened overall availability.
At some point people realized multi-region systems within one cloud platform are prone to failure. They then started deploying their system redundantly to multiple cloud platforms to increase availability. This helped, but created consistency issues. To fix this people started building multi-cloud software systems, creating dependencies across cloud platforms that weakened overall availability.
But no, cloud everything.
Man that's got to suck.
I HAVE NO TOOLS BECAUSE I’VE DESTROYED MY TOOLS WITH MY TOOLS.A very simple example, you do something stupid on a remote machine (either high network usage or CPU usage) over SSH then you can't undo it because SSH becomes unresponsive
Bohr bugs generate an alert and happily meander through normal support channels.
Heisenbugs go through phases -
1. Probation. On continued failure,
2. Restart. If the app or service fails after a restart,
3. Reboot. If the app or service fails after a reboot,
4. Re-image. If the app or service fails after re-imaging,
5. Remove/elimate the node.
The article states:
> The defense in depth philosophy means we have robust backup plans for handling failure of such tools, but use of these backup plans (including engineers travelling to secure facilities designed to withstand the most catastrophic failures, and a reduction in priority of less critical network traffic classes to reduce congestion) added to the time spent debugging.
Another alternative is low bandwidth flag based roll-backs (for instances such as this where the network is congested but not completely lost).
My point being, it seems that modems are becoming less-and-less viable for out-of-band management.
[0] https://scholar.harvard.edu/files/mickens/files/thesaddestmo...
For example GCS had 96% packet loss in us-west. So doesn't it make sense to refund every customer who had any API call to a GCS bucket on us-west during the outage?
Assuming they only refund the service costs for the hours of outage, only the largest of customers will be owed a refund that is greater than the cost of an employee chasing compiling the information requested.
For sake of argument, if you have a monthly bill of 10k (a reasonably sized operation), a 1 day outage will result in a refund of around $300, not a lot of money.
The real loss for a business this ^ size is lost business from a day long outage. Getting a refund to cover the hosting costs is peanuts.
In this outage's case you might be able to argue for a 10% credit on affected services for the month, figuring 3.5 hours down is 99.6% uptime.
but i still agree, it cost us way more in developer time and anxiety than our infra costs, and could have been even worse revenue impacting if we had gcp in that flow
From GCP's top level SLA:
https://cloud.google.com/compute/sla
99.00% - < 99.99% - 10% off your monthly spend 95.00% - < 99.00% - 25% off your monthly spend < 95.00% - 50% off your monthly spend
Probably literally not worth your engineer's time to fill in the form for the refund.
> "[Customer Must Request Financial Credit] In order to receive any of the Financial Credits described above, Customer must notify Google technical support within thirty days from the time Customer becomes eligible to receive a Financial Credit. Failure to comply with this requirement will forfeit Customer’s right to receive a Financial Credit."
FiOS has proactively given me per-day refunds of service without notification on my part. Weird to me that Verizon acts better than Google in this case.
And the part that even Google can’t know, even if they somehow can assemble all of the above: did it matter? Not all servers are created equal.
Small wonder they’re letting customers drive their own reimbursement process.
The icing on the cake was they disapproved one of the ads due to the destination URL not loading.. which was in itself surprising, because everything outside of the affected region was running fine.
They're well aware that ~50% won't bother, so that $10 discount per unit effectively becomes a $5 discount.
Having said that, if Google wants to delight customers, they should give a free tier bonus to all customers for a certain period, but such a thing cannot be fair to everyone.
It'd never happen, "delight" is an Apple principle.
I'd love to see the red binders come down off the shelf, people organize into incident response groups, and watch as a root cause is accurately determined and a fix out in place.
I know it's probably more chaos than art, but I think there would be a lot to learn by seeing it executed well.
The first thing I'll say is that most incident responses are reasonably uneventful and very procedural. You do some initial digging to figure out scope if it's not immediately obvious, make sure service owners have been paged, create incident communication channels (at least a slack room if not a physical war room) and you pull people into it. The majority of the time spent by the incident manager is on internal and external comms to stakeholders, making sure everyone is working on something (and often more importantly that nobody is working on something you don't know about), and generally making sure nobody is blocked.
To be honest, despite the fact that it's more often dealing with complex systems for which there is a higher rate of change and the failure modes are often surprising, the general sentiment in a well-run incident war room resembles black box recordings of pilots during emergencies. Cool, calm, and collected. Everyone in these kinds of orgs tend to quickly learn that panic doesn't help, so people tend to be pretty chill in my experience. I work in finance now in an org with no formally defined incident response process and the difference is pretty stark in the incidents I've been exposed to, generally more chaotic as you describe.
Also during an incident, fingers are never publicly/embarrassingly pointed nor are people blamed. It's all about identifying and fixing the issue as fast as possible, fixing it, and going back to sleep/work/home. For better or worse, incidents become routine so everyone knows exactly what do and that as long as the incident is resolved soon, it's not the end of the world, so no histrionics are required.
The other problem is that it is almost never a single person or teams fault. The reality is that it is everyones fault, and as soon as people accept that they can prevent it in the future.
Lets take a contrived case where I introduce a bug that floods the network with packets and takes down the network. Is it my fault? Sure. But what about pre-deployment testing? What about monitoring - were there no alarms setup to detect high network load? What about automatic circuit breakers that should have taken the machine offline, and instead let a single machine take down the whole system?
The point is that blaming the person who introduced a code bug is lazy, and does nothing to prevent the issue in the future. When a failure like what happened at Google occurs it is an organizational failure, not a single person or team. That is why blaming people is generally not productive.
As mentioned in this thread, it's a lot like listening to air traffic comm chatter.
People say what they know, and only what they know, and clearly identify anything they're unsure about. Informative and clear communication matters more than brilliance.
Most of the traffic is async task identification, dispatch, and then reporting in.
And if anyone is screaming or gets emotional, they should not be in that room.
> And if anyone is screaming or gets emotional, they should not be in that room.
If someone starts yelling around in my incident war room for no reason, they get thrown out. I'm a calm and quiet person, but bugging around during a major incident is one of the few things that make me mad.
I highly advise to read Gene Kranz memoirs "Failure is not an Option" if you work in that kind of environment.
Apparently it was mentioned when the Apollo 13 script writers were gathering stories at NASA, they liked it, and then gave it to the Kranz character.
Who then decided, "Hey, if everyone thinks I said it..." and titled his memoir.
And yes, once we had a workaround in place the customer accepted, we reacted like that. Except we also had our critical-incident whiskey go around. Then the CEO walked in to congratulate us on that project. Whoops. But he's a good sport, so good times. :)
On a slightly different note, "low-level team to have at least one engineer on call at any given time" - this line itself is so true and at the same time it has so many things wrong. Not sure what the best way to put the modern day slavery into words given that I have yet not seen any large org giving day off's for the low-level team engineer just cause they were on call.
There is an understanding of how it impacts your time, your energy and your life that is impressive? To be honest, I feel bad for being so macho about oncall at the org I ran and just having the leads take it all upon ourselves.
Personally I think this is a fair system, and I would hardly call it slavery.
(disclaimer: am Google SRE)
25% time for carrying the pager was to compensate for: 1) Requirement to be able to get to the plant in 30 minutes. Fresh snow? Too bad, no skiing for you this weekend. 2) You must be sober and work-ready when the pager goes off. At a party? Great, but I hope you like cranberry juice.
As the customer who signed the time cards for the pager duty, I thought that was not only fair, but it also drove home to me as a manager that the cost was real and was coming out of my budget, not some general IT budget that someone else took the hit for. This is one case where "You want coverage for your service? Give me a charge code for the overtime." was not just senseless bureaucratic friction, it led to healthier, business-driven, decisions.
It shows in their products (though it's improving)
https://landing.google.com/sre/sre-book/toc/index.html
At my last job, I bought a copy of this book, but we only had the organizational bandwidth to do a few of the things mentioned. At Google, we do all of them.
The incident on Sunday basically played out as described in chapters 13 and 14. There is always the fog of war that exists during an incident, so no, it wasn't always people calmly typing into terminals, but having good structure in place keeps the madness manageable.
Disclosure: I work in Google NetInfra SRE, and while my department was/is heavily involved in this incident, I personally was not.
Also, we're [always] hiring:
https://careers.google.com/jobs/results/?company=Google&comp...
If you're interested in how these sorts of incidents are managed, check out the SRE Book[1] - it has a chapter or two on this and many other related topics.
Disclosure: I work in Google Cloud, but not SRE.
Having been in the ringmasters seat for major incidents ranging from "relatively routine" to "it's all on fire", and had a ringside seat for a cloud provider outage of comparable magnitude to this one - it still fascinates me how creative solutions can get dreamed up under high pressure, and how effective someone to keep the response calm and _feel like it's in control_ is.
I particularly like Gene Kranz "Failure is not an Option". It is more background but it works. In general, it is not crazy hard. You get the roles, you distribute them. Someone can have multiple roles that depends on the size of the incident.
The usual roles i differentiate are Point (think of it as IC if you want), Comms and Logs
Any given production system that works will have capacity needed for normal demand, plus some safety margin. Unused capacity is expensive, so you won't see a very high safety margin. And, in fact, as you pool more and more workloads, it becomes possible to run with smaller safety margins without running into shortages.
These systems will have some capacity to onboard new workloads, let us call it X. They have the sum of all onboarded workloads, let us call that Y. Then there is the demand for the services of Y, call that Z.
As you may imagine, Y is bigger than X, by a lot. And when X falls, the capacity to handle Z falls behind.
So in a disaster recovery scenario, you start with:
* the same demand, possibly increased from retry logic & people mashing F5, of Z
* zero available capacity, Y, and
* only X capacity-increase-throughput.
As it recovers you get thundering herds, slow warmups, systems struggling to find each other and become correctly configured etc etc.
Show me a system that can "instantly" recover from an outage of this magnitude and I will show you a system that's squandering gigabucks and gigawatts on idle capacity.
If it was possible to have this fixed sooner I’m sure they would have done that. That’s not the point of my comment tough.
> From Sunday 2 June, 2019 12:00 until Tuesday 4 June, 2019 11:30, 50% of service configuration push workflows failed ... Since Tuesday 4 June, 2019 11:30, service configuration pushes have been successful, but may take up to one hour to take effect. As a result, requests to new Endpoints services may return 500 errors for up to 1 hour after the configuration push. We expect to return to the expected sub-minute configuration propagation by Friday 7 June 2019.
Though they report most systems returning to normal by ~17:00 PT, I expect that there will still be residual noise and that a lot of customers will have their own local recovery issues.
Edit: I probably sound dismissive, which is not fair of me. I would definitely ask Google to investigate and ideally give you credits to cover the full span of impact on your systems, not just the core outage.
It's hard to tailor a postmortem like this to everyone's individual experience but it is surprising to me that your experience is so different.
Our stack? Multiple OC wan, 10G LAN with 1Gpbs clients. About 4,000+ users, EDU. We are super happy using Google. No complaints! Google is doing great.
Certainly a more interesting story to tell the kids.
Humanity is not only kept safe, but learns about valuable news and offers.
The Bilderburg/Eyes Wide Shut hooded, masked billionaire cultists devised the whole situation as an emergent fitness function. They knew their AI progeny wouldn't be ready to bring the end of days, to rid them of the scourge of burgeoning common humanity, until it could completely outsmart Google DevOps.
_cue music & dramatic squirrel_
Now compare to the free range domain of self-driving cars. If automation fails this drastically, then it does not bode well for self-driving cars.
http://clarkesworldmagazine.com/kritzer_01_15/ https://en.m.wikipedia.org/wiki/Cat_Pictures_Please
No comment. ;)
So the part that sets up the routing tables talking to some global network service went down.
They talk about some of the network topology in this paper: https://ai.google/research/pubs/pub43837
It might be a little dated but it should help with some of the concepts.
Disclosure: I work at Google
Even when they are combined in one device they are often separated on to control plane and data plane modules. Redundant modules are often supported and data plane modules can often continue to forward data based upon the current forwarding table at the time of control plane failure.
Often the control plane module will basically be a general purpose computer on a card running either a vendor specific OS, Linux or FreeBSD. For example Juniper routing engines, the control planes for Juniper routers, run Junos which is a version of FreeBSD on Intel X86 hardware.
That's pretty much the definition of SDN(software defined networking.) The control plane is what programs the data plane - this is also true in traditional vendor routers as well. It sounds like when whatever TTL was on the forwarding tables(data plane) was reached the network outage began.
If there was an analog to a standard kubernetes cluster, I imagine it would be the equivalent of the kube controller manager.
For vmware guys, it would be similar to DRS killing all the vcenter VMs in all datacenters, and then on top of that having a few entire datacenters get rerouted to the remaining ones, which have the same issue.
Does not appear to be true. Tests I was running on cloud functions in europe-west2 saw impact to europe-west2 GCS buckets.
https://medium.com/lightstephq/googles-june-2nd-outage-their...
except for the skateboards, all real sysadmins ride skateboards.
Of course the target audience is probably tiny.
It's Sunday so I guess they are not together. Instead there could be a lot of calls and working on some collaboration platforms. Everyone just staring at the screen, searching, reporting, testing and trying to shrink the problem scope.
If there's a record on everyone there must be a narrator explaining what's going on or audiences would definitely be confused.
It's Google so they have solid logging, analyzing and discovery means. Bad things do happen but they have the power to deal with them.
I suppose less technical firms(Equifax maybe?) encounter similar kind of crysis would be more fun to look at. Everything is a mess because they didn't build enough things to deal with them. And probably non-technical manager demanding precise response, or someone is blaming someone etc.
But in defense, why be admin if you don't look admin?
So it's real skateboard.
Boosted boards are real skateboards too (https://boostedboards.com/) and would make moving through a DC even more effective ;)
SLA CREDITS
If you believe your paid application experienced an SLA violation
as a result of this incident, please populate the SLA credit request:
https://support.google.com/cloud/contact/cloud_platform_slaThis prevents people from pointing the finger at them for not providing SLA credits.
You normally prepare for such a task for a month, and then you hope it will work. In my case (I brought down one the core DNS in Austria for a few minutes, for a very trivial oversight) everyone knew, and after the caches ran out we immediately restored the backup. We weren't on page one in the news as Google.
In the Google case they had no idea of the root cause, so they had to run after this guy who caused it. Only after 4 hours they found him, and they could stop this job. Reminds me a bit of Chernobyl, where nobody told anybody.
I never actually worked in a data center, so keep in mind I don’t know what I’m talking about. Traditional DCs have UPS all over the place, but that will only last a finite amount of time, and your maintenance might take longer than the UPS will last.
What it means to me is that initially some unusually poor decisions were made that triggered an unfortunate and unavoidable events. Very rare is a damage control statement. There is a subtle tone of concern and feeling of blame trough that entire postmortem. This will be buried but if it was investigated thoroughly I wouldn’t be surprised of some serious consequences.
Total speculation. I do not work for google.
Presumably they will get a refund based on SLA for this? Shouldn't they pass that onto their customers?
Does that mean engineers travelling to a (off-site) bunker?
Whenever I post this somebody comes along and says "well that one time us-east-1 went down and everybody was using the generic S3 endpoints so it took everything down". This is true, and the ASG and EBS services in other regions apparently were. BUT, if you invested the time to ensure your application could be multi-region and you hosted on AWS, you would not have seen an outage. Scaling and snapshots might not have worked, but it would not have been the 96.2% packet drop that GCP is showing here and your end users likely would not have noticed.
The articles that track outages at the different cloud vendors really should be pushing this.
GCP was basically even with AWS, and Microsoft was ~6x their downtime according to that article.
> AWS has the most granular reporting, as it shows every service in every region. If an incident occurs that impacts three services, all three of those services would light up red. If those were unavailable for one hour, AWS would record three hours of downtime.
Was this reflected in their bar graph or not?
Also, GCP has had a number of global events, e.g. the inability to modify any load balancer for >3 hours last year, which AWS has NEVER had (unless you count when AWS was the only cloud with one region).
When the primary S3 nodes went down, it caused connectivity issues to S3 buckets globally, and services like RDS, SES, SQS, Load Balancers, etc etc, all relied on getting config information from the "hidden" S3 buckets, thus people couldn't edit load balancers.
(Outage also meant they couldn't update their own status page! [2])
[1]: https://aws.amazon.com/message/41926/ [2]: https://www.theregister.co.uk/2017/03/01/aws_s3_outage/
Source: Im a principal at AWS, historically focused on infrastructure and availability/operations, have been oncall for 20 years, and do some internal incident management as my job.
- the cloud is just a bunch of computers, managed by someone. either you (on-premise private cloud) or by someone else as a SaaS
- building, operating, managing, administering, maintaining a cloud is hard (look at the OpenStack project, it's a "success", but very much a non-competitor, because you still need skilled IT labor, there's no real one-size-fits all, so you need to basically maintain your own fork/setup and components - see eg what Rackspace does)
- it's a big security, scalability and stability problem thrown under the bus of economics (multi-tenant environments are hard to price, hard to secure and hard to scale; shared resources like network bandwidth and storage operations-per-sec make no sense to dedicate, because then you need dedicated resources not shared - which is of course just allocated from a bigger shared pool, but then you have to manage the competing allocations)
Simplify it Google!
Given the fact that the status page was reporting for more than 30 minutes an erroneous infrastructure state and this is google, is it okay for Amazon to put the SRE books into the "Science Fiction" category or should we keep them under tech?
</trolling>
I still feel for the on-call engineers.
Google seems to have created oversight of systems, processes and jobs to be managed by more automation with other systems, processes and jobs.
System A manages its child systems B, which in turn manages its own child systems C and so on. Now the question becomes, who manages the system A and its activities? Automation of the entire tree is as good as the starting node.
Be mindful and make use of automation only of systems that will not be the owner of your business demise. Humans are and should always be the owner of the starting process. Without that governance model, you get google with 5 hours of down time or worst in the near future.
https://news.ycombinator.com/item?id=18428497
login issues:
https://news.ycombinator.com/item?id=19687029
storage system outage:
https://news.ycombinator.com/item?id=19392452
...
So, basically Google created the most unreliable cloud system in the world.
I'm pretty sure that title goes to Azure