Sudden surge of traffic as all their users returns to work?
Hate on the cloud all you want, but AWS has (several flavors of) load balancers and various ways to automatically scale up and down resources (and if you're conservative, you can disable the 'down' part). If you're operating a major SaaS company like Slack and not taking advantage of them, something's gone wrong.
In 2021, how does one keep track of resource starvation at the process, container, os, service, pod, cluster, availability zone and region levels?