While on the topic, we're also eagerly awaiting improved autoscaling (e.g., more responsive, using additional metrics, and scaling down properly). I'd be really curious if you could leverage the more detailed access to instance-level metrics to implement some cool new queue-theoretic modeling: You know roughly how long it takes for an app to launch, you know the current request rate, and you know the time to service requests. You could apply a lightweight Markov model to predict the probability of a given queueing delay in each region within the average launch time and, if so, preemptively launch a new instance before queueing delays even occur. This could be configured to balance a client's tolerance for queueing delays with over-provisioning budget.