For predictive autoscaling, boring old-fashioned forecasting techniques appear to work fine and are very fast and very cheap.
For reactive autoscaling, boring old-fashioned control loops appear to work fine and are very fast and very cheap.
There's still no substitute for understanding your application's design.
If you tried to design and build a factory the way folks propose to rely on autoscalers and ML ("we'll ignore what's in the building and rely on a thin, unintelligible gas of floats!!"), you would be fired.
Examples of hard-coded autoscaling strategies:
Shopping cart or ad delivery network -> anticipate load with extra capacity by scaling up instances before they're needed and kept until they aren't beneficial.
Neighborhood social network:
- scale +1 instances after avg latency > X0 ms for Y0 minutes
- scale -1 after avg latency < X1 ms for Y1 minutes
Adding on this, my current thinking is that it would be possible to better surface this tradeoff by modeling autoscaling as an inventory problem. There's a stockout cost and a holding cost; for a given probability of hitting a cold start and for a given cold start lag, it should be possible to compute the optimal instances on-hand.
Then there are, as you point out, many special business rules. Where I work we have a pool system for testing environments. It works reasonably well until a particular team, sitting upstream of approximately half the company, makes a release. Then suddenly everyone's pipelines kick off simultaneously. So now they have a system which watches for signs of impending release and pre-emptively scales up.
Autoscaling should not be your first line or final line of defence. It is a powerful tool, not a magic wand.