I'm curious, as I might learn something from interesting asnwers to the above questions.
I'm curious, as I might learn something from interesting asnwers to the above questions.
Yes, it absolutely makes sense. You never know when a typo in a config file or bug in your job-management code will go bonkers and try to take over all resources. Same think as disk quotas on computers with tons of users: you want to limit the damage that a user can accidentally (or not...) do.
Good quota systems saved my team's bacon way back when a few times, when fewer people were at the company; I can only imagine how useful they are at what, 10-20x the size?
The SLA would have to be something like "we guarantee 99.999% write availability unless team A or team B or team C makes a mistake", so now you have to look up what team A and B and C are up to and what do they guarantee. Instead of the more sensible "we guarantee 99.999% write availability unless you are out of quota".
Quota systems for the win -- if only as a shared agreement on who can do how much damage / take how much resources before they're automatically stopped and more resources must be justified through some review process (likely involving budgetary concerns).
Source: worked at a FAANG for a decade, saw many big incidents. Almost all were the confluence of several smaller issues that, had they happened alone, wouldn't have been newsworthy.
But my recollection is that most SRE teams would have had exactly such things -- either as really crazy borgcfg (sorry, I mean kubectl) code, or as a borgmon (sorry, prometheus) alert.
The former were a real PITA to work with (borgcfg hadn't been designed with that in mind), and the latter were only reactive (and thus unable to warn about things until they were fast becoming a real problem).
So most likely those safeguards were somehow disabled or made ineffective by the exact circumstances of the problem, possibly the speed at which things developed.
In a small org where you know/roughly-know who is doing stuff then it is probably less of a concern as you can go walk over to their desk and have a chat about what they are doing or let them know they're sending too much traffic etc. This is harder to do with an org with 10s of thousands of engineers spread across the globe and in different timezones - you might suddenly get huge amounts of traffic at 3am from a team on the other side of the planet who just launched a feature you had zero idea about, and they're hammering you with 100x the QPS you usually get.
By everyone else having quotas on how often I can call their systems, they protect every other client they have. They protect customers who are using anything other than my website. Otherwise, someone could take down all of AWS by spamming just one system. I want them to throttle me, quota me, because I want to protect other customers if I get compromised.
And even if it's one internal system calling another, it's better to have protection at every layer than to try to only have it on the outer layer.
This can burn us, of course. We're always growing, and quotas need to be managed (automatically or manually). It all has to be done carefully.
Edit: as always, I speak for myself and not for AWS in any official capacity. I'm mabbo, not jeffbar :)
I imagine once they knew what the issue was, they would have been able to fix it as i'm pretty sure that service doesn't need to be authorised by itself to be configured ;)
The trouble I imagine was getting visiblity of what the issue was...
You'd be surprised at the circular dependencies that creep up at Google scale...
(I used to work there and we had last resort secure access mechanisms, but I wouldn't be surprised if the normal front door wound up shut pretty hard for this one)