Hmm I would like to hear more about these problems that prevent people from spinning up instances. Is it a frequently occurring problem or only happens rarely (e.g. when APIs are down). Also are they managing the instances themselves or using EC2's AutoScaling cluster? I run a dynamically scaling cluster on EC2 and have not run into the problems like you mentioned, so I would like to hear more about them if possible. Maybe they are spinning up and down too rapidly and exceeds the API rate limit? You'd have to have a bugged/bad provisioning system to accomplish that though...