Most of the situations where we needed to drastically scale up were known ahead of time as well (e.g. campaign from customer), and we would preallocate instances or even more clusters.
I may be forcing my memory, but if I'm not mistaken, our auto scaling was setup in a way that the system could handle sudden load increases of ~50% without noticeable disruption. Spikes bigger than this could lead to increased latency and/or error rate.
That said, it's a trade off between efficiency and load spike tolerance. I trust that the trade off is made with informed decision.
Unless you co-mingle online and offline (batch) traffic on same hosts, flat response times and high utilization aren’t compatible.
I don't think that relatively low utilization rates is the scenario that requires "informed decision". The only tradeoff in low utilization rate scenarios is cost, which might be outright cheaper and irrelevant once you do the math on the tradeoffs of using reserved instances vs the cost of scaling up with on-demand instances.
You need to make a damn good case to chronically underprovision your system and expect it to autoscale your way into nickle-and-dime savings.
Worth noting that requiring teams to use 3 AZs is a good idea because you get "n" shaped patterns instead of mirror shaped patterns, which have very different characteristics for resilience and continuity.
AWS does that with their lambda arch to reduce waste.