My biggest issue with serverless is the high cost at scale, which could be mitigated with this type of architecture. Are there any other benefits or concerns regarding something like this?
My biggest issue with serverless is the high cost at scale, which could be mitigated with this type of architecture. Are there any other benefits or concerns regarding something like this?
If that's not fast enough, you can bake an AMI and scale more proactively (e.g. scale when your server is 70% utilized rather than 95%)
EC2 autoscaling takes a minute or two. This means you need to run with enough headroom to absorb spikes before you're able to spin up more resources. Depending on traffic patterns this can get very expensive (or not a problem at all, with predictable loads).
Our customers are concentrated around N. America, with minimal activity at night and weekends. We have a complex product with a wide surface area. Some of our services don't see usage constantly throughout the day.
Our usage is unpredictable and spiky, worsening our ability to cost optimize EC2 autoscaling.
Then there's staging environments...
Smallest EC2 instance will fall over with a sudden traffic spike, and requests will fail for the couple minutes it takes for autoscaling to complete.
Additionally for HA, a machine must be able to die without exposing us to the risk of a traffic spike taking us down.
For example, your base application runs with t2.micro. As soon as traffic starts going up, you launch additional medium/large instances. The load balancer would have to understand this also, and do weighted routing (ideally based on app server loads) so the first micro instance doesn't handle the same volume of traffic as the rest.
I've never seen a setup like this before, and frankly, part of me wonders whether the savings would be worth the effort of building it in the first place (vs just running with medium instances, or just scaling to dozens of micros to handle the load, or just using serverless).