For example, talking about the CPU utilization of a web service, your service can perform a mix of time-critical work (serving queries) and background tasks (indexing, etc). If a server is running at 100% CPU but 30% of that is spent performing background work, the amount of slack available for a sudden surge in demand is 30%, and I suspect it's almost never the case that a large service sees an unexpected 30% load increase in less time than it takes to boot another machine, so such a system could be both efficient and robust with few downsides.
Implementing that system isn't as easy as just throwing more Kubernetes nodes at your problem, especially given that the ecosystem of tooling isn't designed to make it easy. For example, it would be really nice if load balancers used realtime performance metrics to balance traffic at a millisecond level.
Perhaps the real lesson is that "many of the easiest ways to achieve robustness involve sacrificing large amounts of performance", but I reject the idea that we should use that as an excuse to accept terrible performance, with all the monetary and environmental impacts it brings.