Autoscaling, Welcome to Google Compute Engine
googlecloudplatform.blogspot.com
googlecloudplatform.blogspot.com
It wasn't always this way -- at one point we were having so many issues with AppEngine's autoscaling that we ended up in a meeting across from Ben Traynor and heard some of their engineers discussing post-mortems of scaling failures in gruesome detail. They've come a long way since then and AppEngine has really found its way from something of an internal science experiment to what I think is the best competitor for AWS' offerings.
FWIW, the "backend" or "basic" scaling module that was available for both AppEngine and GCE was so poor and unreliable that it made it nearly impossible to scale anything in a way other than manual. I honestly hope they deprecate this and roll the few features into the autoscaler.
The entire scaling infrastructure of AppEngine was horribly broken and unreliable up until a year-and-half or so ago, but they've put a lot of work into fixing that recently and it's in awesome shape now.
It isn't the compute engines job to determine if workload is malicious. That is a separate but real problem.
Ultimately though, from an instance's perspective, there's not a lot of difference between a malicious DDoS and a genuine traffic spike. You can mitigate this (a) by setting low max-instances thresholds when you aren't expecting high traffic, and (b) upstream filtering or QOS management to filter our DDoS traffic from legitimate traffic.
Even if you could identify malicious usage, it's still a matter of policy on how you want to respond to it - given that if it can reach your VM it will likely be affecting legitimate traffic too.
Say I expect high traffic only during mornings, and not want to handle it during the evening if it occurs. Any chance this is in the making?
You can do auto-scaling without Amazon's secret sauce (or Google's), it just relies on triggers to spin up new boxes. I'm not very familiar with the Google setup, but you could use the Amazon API to pull metrics about boxes, then send a request to add another box to the auto-scaling group manually if you wanted.
> How do you prevent oscillations in resource allocations when the traffic is fluctuating?
Oscillation is kind of the point -- if you need x servers to serve y hits per hour, you'll have x servers when you need it. But when you're down to y/2 hits per hour, you don't need to be paying for x servers anymore. It just depends on load -- for example, you can set the group to add a box when average CPU rises above 70 percent, and then set it to remove a box when average CPU falls below 30 percent.
> How do you know if e.g. an adversary is just DDoS'ing your application to make it more expensive to run your application? I guess there must be some kind of threshold that you can set? But then this would cut off the website for the real customers...
This isn't a problem for auto-scaling, it's a problem for web stacks – the whole point of the "auto" part of it is that increase traffic means increased resources. Going counter to that gets complicated quickly. I'm sure there are folks who know more about DDoS mitigation than I do, but putting protection in front of your app with something like Akamai should help, as can keeping an eye out for bizarre spikes from certain IP ranges.
- reallocation from pools before scaling out, which is not very complex either, but it adds a bit of complexity, especially if you optimise for cost along the way.
- linked scaling requirements, e.g. a team grows but there is a bottom and top limit to team size and when you get a new team you also need a supervisor
On oscillations, you can weather them out by the way you treat your driver sampling window (static or variable) and a cost function. Normally oscillation is not bad for the resource allocation scenario, it is the intention of it, but for FTEs the cost can be more difficult to manage than for computational resources.
And based on the comments, has had it for 1-2 years now.