Designing a scalable API on AWS spot instances
blog.adapty.io
blog.adapty.io
What's fantastic about the autospotting implementation is:
1. It dynamically replaces existing instances with spot instances just by setting a tag
2. Rather than replacing the instance with a fixed spot instance type, it will choose the cheapest that fits the requirements.
3. If there are no spot instances available that fit the requirements, it will spin up on demand instances until spot instances are available.
If you have experience working with spot then you will know that these are really outstanding features that hopefully amazon will bake-in in the future.
I'm also missing a discussion about designing for interruption, either by not keeping state, or by being able to shed state quickly, to be picked up by other instances.
Also, if you set up EC2 spot with a launch template or ASG with very differently-sized instance types (to reduce risk of running out), is there a way to even out the load coming through an ALB? The least-connections scheduling can help in some cases, but a connection might not map 1:1 to one unit of load. The ALB can use weighted balancing, but on the target group level. Dunno how easy it would be to allocate different instance sizes to different target groups and weigh them accordingly.
We have this setup with two capacity providers (FARGATE_SPOT and FARGATE) with a 75/25% split, meaning that even if there are no spot instances available we will still be up.
The benefit of Fargate being that we don't need to care if certain instance sizes are not available as that is handled by AWS.
Fargate Spot is about a third of the price of Fargate (at least in eu-west-1 now according to: https://aws.amazon.com/fargate/pricing/ ); so the savings are roughly identical.
Re risk of running out, our current strategy is to use different-but-closely-similar instance groups; so for example we have an autoscaling group running a mix of:
- m5.large - m5dn.large - m5n.large - m5ad.large - m5d.large
Which are the same price on Spot instances, but I'd wager it'd be pretty rare to have all these families reclaimed at once.
(We also use some on-demand only ASGs with lower priority in the cluster-autoscaler to ensure that if it _does_ happen, then we'll have a fallback)
In reality, you would like to compare Reserved Instances as you can get 60% discount.
So in us-east-1:
- spot costs: 35-43% of on-demand
- RI 1 year standard: 60% of on-demand
- RI 3 year convertible: 46% of on-demand
So if you have some base load that you can commit to running for 3 years, the price gets often at spot range while not having to worry about losing capacity.
In-reality sometimes combining reserved for some base capacity 60% + 40% spot for spiky seems to be the winning combination for many companies.
Even if in practice AWS never sees large spikes in compute demand and corresponding large scale instance preemption, most businesses I've worked with won't accept the risk of having OLTP systems be taken down at any time.
No longer does spot seem to be a service where one can get a bargain for their compute intensive offline/batch workloads that are much more tolerant of preemption.
Given that the spot prices seem to be very flat, and preemption is rare, amazon presumably have a fair bit of underutilized capacity, does anyone know if amazon uses this capacity themselves, or offers more aggressive spot pricing to select clients?
Some tangential thoughts:
* Is there an AWS API that returns the cheapest availability zone in a region for a given instance type? Or is the GUI that's screenshot in the blog the only way to see?
* I have seen the 90%+ cost savings for certain instance types
* Sometimes you lose a spot instance, look at the pricing history graph to confirm the price spike, and don't see any spike that was above your bid price... it can be frustrating
Having a bid above spot price does not guarantee you'll keep the spot instance. AWS can terminate a spot instance at any time if they need the capacity—That's the deal. It used to be more closely tied your bid price, but they've been moving away from that.
This is a recurring topic here on HN and it boggles me and makes me wonder if people know that there are other platforms than aws, azure and google cloud out there that are very capable and much much cheaper.
Unless any of the big 3 has a feature or certification you need I don't see any reason to use them at all due to the insane complexity and cost.
So why do you or your company who uses any of the big 3 use them if you had to cut cost at some point?
Lots of providers claim to offer object storage, but try hitting them from couple thousand cores and they all tend to immediately fall over.
Unless that limitation has changed in the last couple of years; I can easily make a system beat that, if that's the requirement.. and ultimately it does come down to understanding requirements. :\
I think generally people forget that cloud is just computers too, it's really not anything special, and amazon/google/microsoft are solving the general case (and, doing so well, actually) but this comes at a high premium.
[1]: https://aws.amazon.com/premiumsupport/knowledge-center/s3-ma...
https://www.digitalocean.com/docs/release-notes/upcoming/spa...
The main reason we wanted to work with DO is that the egress traffic cost is reasonable, unlike AWS or Google Cloud where it is questionably high.
Honestly it soured my opinion on DO a bit. I'd previously had a droplet there for several years, which never seemed to experience any issues at all.
I'm not sure where the idea came from that this is hard to do.
To answer your questions in order:
1) People warned aggressively about lock-in of cloud providers with proprietary extensions and you must have chosen not to listen; so, I have little sympathy.
I'm not saying it's black and white; but I hope that you got your velocity required to hit market faster and have made more than you spent because this is the price you paid; and now to get out you'll have to invest a little time on cleaning house, that's the reality of lock-in.
2) S3 is pretty easy as there are "s3 compatible" FOSS projects; min.io comes to mind, or ceph with a RADOS plugin, or Riak with the s3 plugin... there's also s3proxy with a multitude of backends..
3) Aurora serverless is replaced by knative
4) Beanstalk is just classical servers with an auto-scaler component, auto-scaling depends greatly on your provider, so I can't say how easy or hard it will be, if you're using kubernetes then understanding your load should be easy at least.
I don't understand the argument: "I want to outsource understanding but I want to save costs";
You can think of things as a spider-graph of three points:
Quality -- Low-Cost -- Knowledge-Required.
No solution can score high points on all three reasonably.
Anyway, things are a bit skewed because I'm an infrastructure type, and people in my profession really do think of systems administration tasks as being "very easy" and if done right soak up nearly no time at all, but developers don't like hearing that because sysadmins are "old world".
I don't really care if you're paying someone elses sysadmins or not, the fact remains that you're going to be spending something in that area, and if you balk at the cost of cloud then maybe taking ownership of what they do can help optimise costs.
Obviously they put a premium on their own time in these areas.
Full Disclosure: I work at AWS building tools to help customers do cost optimization.
----
> Have you actually done a cost-benefit analysis on some of these solutions?
Yes, I even gave a talk at google in stockholm about it.
For my use-case, hybrid was best, with no cloud lock-in aside from Google Storage Buckets (which can be replaced) but I went into detail about that in the talk.
> Take your Riak / S3 plugin. What do your servers cost to run that cluster?
Depends a lot, don't you think?
> How much time do you spend managing it?
Depends again, if it's anything like my elasticseach clusters then about 2-man hrs/mo.
> How do you test your backups?
Continuously, and with alerting.. and, you should be doing this anyway.
> Are you going to target the same SLOs for durability that S3 offers?
Depends on the business, the whole point of SLO is that you pay in what it's worth to the business.
> Do you run multi-data center for high availability?
Depends on SLO.
> In many cases the cheaper or self-hosted solutions have costs that you aren't accounting for.
Yes, physical machines often need some hand-holding, VPS's can have brown outs, but this is true in AWS's EC2 anyway.
Ultimately, this is where the cost increase will be.. but defining it is important, I've deployed cloud and physical (as stated) and it's true that physical machines are not as problem-free as our GCE ones- but we pay about 50% less than the GCE equiv instances, so it's "worth" spending time automating the unpredictable.
> Sometimes that's fine, but "just run it yourself" is as worthless as saying "just ship it to AWS" unless you actually think through the impact.
This is kind of the main point I always make.. understand your trade-offs, don't buy into proprietary tech. Cloud is a fantastic way to prototype and bootstrap but it's /usually/ better to have a migration plan to optimise costs in the future.
If you fail to take that into account then I don't have sympathy for you, because you put the project at risk. Financial in-viability is a risk.
> understand your trade-offs, don't buy into proprietary tech. Cloud is a fantastic way to prototype and bootstrap but it's /usually/ better to have a migration plan to optimise costs in the future.
Preach.
I think that's a myth. People assume it's true because it should be true. I don't think it is true.
> 2nd AWS is not only a VPS provider.
It's a glorified VPS provider. Most of the stuff doesn't matter to most of the people using AWS, but they go for it so they can put it on their resume and because they don't want to get fired for choosing something that's not a big name.
Load balancing, fault tolerance, high availability, arbitrary scale, messaging, orchestration, autoscaling, warehousing, big data processing, identity management, desktop management, secrets management, container registries, source code management, build tools, hardware test suites, gpu hardware, observability tools...
Those of use that use cloud providers know full well why we use them (and certainly know when not to).
Re #2: I love all the other stuff. I never use EC2. Lambda, Cognito, DynamoDB, S3, CloudFront, Route 53, and API Gateway are my default stack, managed by Cloudformation. Granted, I'm doing smaller projects, but the costs are minimal, the setup time is trivial, the documentation is excellent, everything is nicely compartmentalized and 'just works' together. And I only pay for the actual traffic to the site.
It makes sense that bargain-bin providers would offer inferior reliability, but there are providers out there other than the big 3 cloud providers and the bargain-bin VPS providers.
GitHub for instance is apparently [0] hosted by Carpathia [1].
[0] https://github.com/holman/ama/issues/553
[1] http://www.carpathia.com/ (they should really fix https://www.carpathia.com )
In that period (2015-2018) I used to run a fairly well known French website on OVH, and their network was very unstable, from equipment failures to way too many fat-fingering of routes. If you're able to easily switch traffic between OVH and AWS you're in a far better position than most people.
But at the end of the day outages happen everywhere, including AWS. We also have some kit at Hetzner and I think that a redundant setup across OVH and Hetzner will be a fraction of a cost of single AZ setup in AWS and yield far greater uptime.
We’ve commoditized the servers and services (cattle vs pets)...why would we treat the providers any differently? Use cheap components and lots of redundancy.
Because it takes X hours of effort to cut costs and using another provider would have required Y hours of effort where Y >> X. Y may be greater due to reliability issues, missing features that you need to build yourself, training costs of new employees, etc.
edit: Also X is paid once you're succeeding, Y has to be paid before you're succeeding which makes Y even more expensive in opportunity cost.
If you're really looking to save on costs, hosting solutions (i.e. hosting racks or some managed solution) are probably what you would need to look at. But then there are other costs involved there as well (infrastructure team, upfront capex costs etc.), but it might still be worth it if you run e.g. a ton of batch processing over a ton of data.
If someone is interested in learning all the AWS concepts, here's an awesome e-book which is written by the legend Daniel Vassallo himself.