Preliminary Incident Statement
status.cloud.google.com
status.cloud.google.com
I'm curious, as I might learn something from interesting asnwers to the above questions.
I imagine once they knew what the issue was, they would have been able to fix it as i'm pretty sure that service doesn't need to be authorised by itself to be configured ;)
The trouble I imagine was getting visiblity of what the issue was...
You'd be surprised at the circular dependencies that creep up at Google scale...
(I used to work there and we had last resort secure access mechanisms, but I wouldn't be surprised if the normal front door wound up shut pretty hard for this one)
Yes, it absolutely makes sense. You never know when a typo in a config file or bug in your job-management code will go bonkers and try to take over all resources. Same think as disk quotas on computers with tons of users: you want to limit the damage that a user can accidentally (or not...) do.
Good quota systems saved my team's bacon way back when a few times, when fewer people were at the company; I can only imagine how useful they are at what, 10-20x the size?
The SLA would have to be something like "we guarantee 99.999% write availability unless team A or team B or team C makes a mistake", so now you have to look up what team A and B and C are up to and what do they guarantee. Instead of the more sensible "we guarantee 99.999% write availability unless you are out of quota".
Quota systems for the win -- if only as a shared agreement on who can do how much damage / take how much resources before they're automatically stopped and more resources must be justified through some review process (likely involving budgetary concerns).
Source: worked at a FAANG for a decade, saw many big incidents. Almost all were the confluence of several smaller issues that, had they happened alone, wouldn't have been newsworthy.
But my recollection is that most SRE teams would have had exactly such things -- either as really crazy borgcfg (sorry, I mean kubectl) code, or as a borgmon (sorry, prometheus) alert.
The former were a real PITA to work with (borgcfg hadn't been designed with that in mind), and the latter were only reactive (and thus unable to warn about things until they were fast becoming a real problem).
So most likely those safeguards were somehow disabled or made ineffective by the exact circumstances of the problem, possibly the speed at which things developed.
In a small org where you know/roughly-know who is doing stuff then it is probably less of a concern as you can go walk over to their desk and have a chat about what they are doing or let them know they're sending too much traffic etc. This is harder to do with an org with 10s of thousands of engineers spread across the globe and in different timezones - you might suddenly get huge amounts of traffic at 3am from a team on the other side of the planet who just launched a feature you had zero idea about, and they're hammering you with 100x the QPS you usually get.
By everyone else having quotas on how often I can call their systems, they protect every other client they have. They protect customers who are using anything other than my website. Otherwise, someone could take down all of AWS by spamming just one system. I want them to throttle me, quota me, because I want to protect other customers if I get compromised.
And even if it's one internal system calling another, it's better to have protection at every layer than to try to only have it on the outer layer.
This can burn us, of course. We're always growing, and quotas need to be managed (automatically or manually). It all has to be done carefully.
Edit: as always, I speak for myself and not for AWS in any official capacity. I'm mabbo, not jeffbar :)
Talk: https://eventsonair.withgoogle.com/events/autopilot-research...
The paper and the system that caused this incident are different. Google has a ton of different automated systems for maintaining production.
The paper you linked is for a system called autopilot, which can scale a specific jobs memory and CPU usage up or down depending on historical load of that job. Think of it as having different instance sizes on GCP or AWS, then you have a tool that monitors how much actual CPU and memory your job uses on the instance, in will upsize or downsize your machine depending on load over time.
The system that broke based on the incident report above had to do with quota that a given production system could use as a maximum. Similar resource management automation, but very different systems.
So one is about fine-tuning, while the other is about preventing excessive resource usage.
At least it's not just us pleb users at the mercy of their automated systems. I wonder if someone internally appealed the decision and got told their review had been carefully considered and subsequently declined.
"But picking up the phone doesn't scale"
Me: looks at size of fines paid in last 10 years.
> Many of our internal users and tools experienced similar errors, which added delays to our outage external communication.
Also:
> We will publish an analysis of this incident once we have completed our internal investigation.
Since "Preliminary Incident Statement" is basically a subtitle in the OP, I've put that in the title field above.
Editorializing would be changing it to something like “Breaking the silence after a catastrophic outage caused by our gross incompetence”.
Or what was that one we had the other day?
"The death of Google"
.... so much for limiting Blast Radius
A good infra architecture has the “blast radius” of any issue confined to only part of your infra fleet. Avoiding the global outage.
Think of it like a navy ship. When a mussel breaches the hull, the ship is designed to contain the leak in a single section. Avoiding the entire ship sinking. Similar pattern is desired in your service and compute infrastructure.
The outage google just faced is equivalent to a single missel taking down an entire navy fleet.
> The outage google just faced is equivalent to a single missel taking down an entire navy fleet.
I really thought you were going to say "is equivalent to a single missile taking down an entire ship" - but your extreme analogy made even more sense in this context.
I knew barnacles were a problem, I hadn't realised there were other dangerous bivalves. :)
As the time of the outage increases, service health probably trends to zero without being able to manage the service... but for a few hours, it's not a disaster. Usually.
The time when Google was special is long gone.
just kidding.