And the postmortem is basically like "yep can't really blame anyone for not knowing this weird ass thing or unfound bug".
Software, man.
And the remediation procedure is quite simple too - scale down. If the fleet can't handle it, increase your throttling. If even with increased throttling you can't handle the load and everybody's timing out (but you still can't scale out), you start spinning up new load balancers, and applying weights in DNS where at least some percentage of your callers are getting through while others see full failure.
The key point is in anything related to such an outage, even if it takes a long time to recover, your customers shouldn't see a full outage.
So it was either an entirely different class of problems or Roblex has a huge ops skills gap. I'm betting the former.
edit: never mind, maybe it's the latter after all: https://twitter.com/NIDeveloper/status/1454773313792880640
1) There are a lot of dimensions
I think it is fairly common to focus on a throughput-related dimension like requests per second and test that very thoroughly, while not paying enough attention to other aspects. We have services where, if we were told requests per second were going to go up 100x tomorrow, I wouldn't be too concerned, but if we were told request size were going up 2x, I'd be freaking out.
Even when you think you've identified all of the dimensions, there are usually ones you missed and/or they interact in weird and unexpected ways.
2) Performance can appear to be linear when it really isn't
So many outages I've seen have resulted from everything being fine until some threshold is hit, at which point it doesn't start to slowly degrade, but instead immediately explodes. Often times due to feedback loops (GC activity in garbage-collected languages often can behave this way) or because some cache overflowed (data that is queried at a high rate no longer fitting in memory on a DBMS comes to mind).
Because the business had been around for a while, there was two separate ways to deploy and manage apps: old and new.
We reconfigured how code was deployed within and between racks for increase resistance to racks or even whole colos dying. This change was made in the new deployment system. And all was well for most of a year. Software and services gradually migrated from old to new as people had the time and inclination.
The new deployment also packed services more tightly, so most computers kept copies of most binaries on them.
Then we crossed one of those tipping points. On a code update, several thousand servers across the world demanded a small set of binaries from a distributed data node via the newer deployment method. Which would be fine, except we had degraded service from a couple TOR routers due to a separate bug. And so the distributed data nodes hosting this particular set of binaries were less "distributed" and more "just one poor about to be overloaded machine and network."
So a few thousand machines demanded a couple hundred megs nearly simultaneously. And when they didn't get it in a timely fashion, all triggered a rollback. Simultaneously. Which demanded another couple hundred mb of code. Simultaneously. And then the freakouts started.
And because other deploys were in flight, there was more than one service deploy ongoing. And they all started fighting each other for bandwidth.