Leveraging mispriced AWS spot instances
pauley.me
pauley.me
Kudos to the author on producing this.
Quick, small fix: This instance is shown as having 0 GBs of memory, but in fact it has 0.5 GB. https://instances.vantage.sh/aws/ec2/t4g.nano
Do you mind opening an issue on the repo here? https://github.com/vantage-sh/ec2instances.info
Thank you for the report!
We see similar mismatched pricing all the time and take advantage of it. One additional area not called out here is the difference between c5.24xlarge and c5.metal instance pricing. These are pretty much identical hardware but metal instances are often cheaper.
As you go down this path, do expect to see a lot of weird things that you'll have to track down. For example, when we introduced metal instances we found that the default ubuntu AMI launched with a powersave cpu governor. Non-metal instances don't support CPU throttling so it never came up with c5.24xlarges. When we first launched metal instances the performance per instance was significantly worse and took a bit of work to track down.
Recently we've seen a lot more spot interruptions and it's pushing us to incorporate more 6th gen instances to get us more diversity. We've also temporarily switched to capacity optimized over price optimized and we've enable capacity rebalancing.
It's absolutely a win for us from a pricing perspective. Our traffic is extremely variable each day and very seasonal throughout the year. RIs don't make sense given <12 hrs daily peak and 10x difference between July and September. However, just plan for some odd surprises along the way.
Caveat operator.
(I’m sure parent commenter is either not exposed to this scenario or has otherwise mitigated against it)
We also have a second, on demand, ASG ready to fire up at a moments notice if something were to happen with capacity.
We also heavily leverage managed services for state.
They spot an inconsistency between two prices, and decide that the fair market value must be the very highest part of the spread. Anything under this is therefore "under-priced".
Is it not possible instead that people are overpaying for the popular ones through sub-optimal bids - instead of simply assuming that only these inflexible/least sophisticated bidding strategies represent the fair market value.
They actually go further. They assume that AWS could realize this value, and that encouraging more flexible bids through tooling etc. would move everyone to the top of the spread, instead of smoothing it out towards the average. And that what is essentially a price-increase can be achieved without hurting the overall value (price-performance vs. flexibility). Given the entire point of this is auctioning unused cycles at a discount, clearly any overall increase would decrease the overall demand.
Having said this it's a great article. I think the overall quality of the article made it so surprising to see this missed.
Keep up the good work
Many workloads do not have deployment configuration supporting a non-homogenous fleet of instances. Over time this will be addressed, but it could be a current major contributor to the discrepancies viewed.
Instances with attached NVMe are available in much lower volumes than others, as are AMD instances. Obviously these pools cannot be used as a drop-in replacement for non-"d" instances or Intel families.
That being said, the substitute instances considered could be trivially accepted by any task running on the original instance, so long as it doesn't misbehave when given too many resources. In the case of vCPU, you can even hide extra vCPU cores, so a c6g.xlarge can be made effectively indistinguishable from a m6g.2xlarge by disabling the vCPUs at the hypervisor level.
You are hypothesizing that the price differences produce "lost" revenues.
An alternative hypothesis can be that the price differences produce similar or higher level of revenues for AWS through price segmentation, with Amazon recognizing the lack of adoption of certain spot instance bidding features and auction markets reacting appropriately.
Unless you have the capacity and quantity demanded for each instance types, you can't prove your hypothesis. You are assuming scenario 3 (below) with no insights into price elasticity of the underlying customers.
Example:
Baseline:
Instance types A and B are equivalent.A is priced at $3, with capacity of 1000, quantity demanded of 800. B is priced at $2, with capacity of 1000, quantity demanded of 200. Total quantity demanded = 1,000.
Revenues from instance type A = $3 x 800 = $2,400 Revenues from instance type B = $2 x 200 = $ 400
Total revenues = $2,800
Scenario 1: All customers purchase instance B instead due to better price discovery.
Revenues from instance type A = $3 x 0 = $0
Revenues from instance type B = $2 x 1,000 = $ $2,000
Total quantity demanded = 1,000.Total revenues = $2,000
Amazon loses $800 in revenues, there are no "lost" revenues" recovered.
Scenario 2: Amazon changes instance type B price to $3. Total quantity demand decreases to 900 due to price elasticity of instance type B customers.
Revenues from instance type A = $3 x 800 = $2,400
Revenues from instance type B = $3 x 100 = $300Total revenues = $2,700
Amazon loses $100 in revenues, there are no "lost" revenues recovered.
Scenario 3: Amazon changes instance type B price to $3. Total quantity demand remains at 1,000.
Revenues from instance type A = $3 x 800 = $2,400
Revenues from instance type B = $3 x 200 = $600Total revenues = $3,000
Amazon recovers $200 in "lost" revenues.
Assuming the market is in equilibrium, the above scenarious aren't realistic, as demand at the market price would equal supply at the current price (roughly, of course).
Suppose there are 1000 c6g and 200 c6gd, with equilibrium price of $3 and $2, respectively (i.e., all instances have demand). Amazon re-SKUs c6gd as c6g until there are 1100 c6g selling fro $2.90 and 100 c6gd selling at $2.90. Total revenue is $3480 vs. $3400. Of course it's impossible to know the true numbers without hidden knowledge of the market, but this is more akin to what would occur. Amazon effectively has a risk-free arbitrage opportunity here, so it stands to reason that there is revenue to be made. Customers don't have this option (since you can't short spot instances), so the best you can do is diversify and save money.
Edit: Actually, the AWS spot market is often out of equilibrium in a way that makes this reselling even more effective. For instance, in the example in the article the c6gd instance is actually pegged at the minimum price, so some number of those instances could be resold as c6g without moving the c6gd price at all.
Spot instance capacities are a function of the all instance capacity for the same type and on-demand instance usage. Spot instance pricing can influence the quantity demanded of on-demand instances of the same type, and vice-versa.
Anyhow, there’s no way we can figure out whether you’re right or wrong with any reasonable level of certainty.
Straightforwardly: All hosts with space for a c6gd spot instance have space for a c6g instance. If Amazon is willing to host a c6gd instance in that slot for $X, they should be willing to also host a c6g instance there for $X.
In financial markets, the way this gets handled is through arbitrage: someone will buy the equivalent of the c6gd instance, and sell the c6g part for the higher price (they may also sell the "d" part for even more money). This has the effect of "correcting" the price. The AWS spot market does not allow you to do arbitrage, and AWS doesn't appear to do the arbitrage for you.
AWS probably likes this inefficiency in their market: some instance types are more popular than others, and some customers make assumptions that require them to use a very specific instance type (ie a c6gd would not work as a substitute for their c6g instance). However, the vast majority of users probably could work just fine if their c6g instance were a c6gd, and don't look for the arbitrage opportunity. That means Amazon gets paid extra.
The reality is that direct c6gd demand might be an order of magnitude lower than c6g direct demand - if AWS can get some more flexible people to adopt c6gd by offering a lower price, c6g capacity is slightly stabilized for on-demand usage by people who don't value the flexibility.
Also note that c6g to c6gd has a non-zero switching cost - extra NVMe on the instance adds a new source of potential hardware failure, increasing the probability of termination very slightly. There might be other software-related costs depending on whether your application makes any ill-advised assumptions about attached storage during setup.
So overall, I would just be happier to read this article if it was framed as "PSA: having more features in an ec2 instance is sometimes cheaper! Don't rule yourself out of extra savings by making overly-constrained fleet requests." The extra commentary about foregone revenue makes too many assumptions and detracts from the core point.
The fact that you have to host a c6gd to get that price instead of a c6g is an inefficiency in the spot market that likely makes Amazon money, but is a little customer-hostile. I think the article is probably wrong that Amazon is foregoing revenue due to this. This is a form of price discrimination and it is likely making Amazon money, but in a scummy way.
Also, it really helps to analyze at the AZ level. Certain AZs lack instances or have very low spot availability and contrary to recommended best practice, reducing AZs can sometimes be beneficial (I am looking at you eu-central-1a).
While lowest price sounds nice, they can be really messy in terms of spot interruption rate. It is much better to set a max price and choose capacity optimized with as many instances as possible.
FYI, AZ names are not universal. Your eu-central-1a might be someone else's eu-central-1b.
Simplify do:
curl 'https://ec2.shop?region=us-west-2&filter=m5&json' | jq
You can pipe it to whatever your system store to get the real time price without dealing with AWS Price API
>c6g.2xlarge→c6g.4xlarge→m6g.4xlarge→r6g.4xlarge→r6gd.4xlarge
A long standing ticket in my personal project backlog is comparing different instance types performance. I'm not sure this equivalente is without caveats.
Anyhow, the reason "misprices" exist is because:
- Many AWS products are elastic but only allow one to choose a single instance type. So you need to guess the best instance for a workload and stick with it.
- No AWS product exposes a "Just give me the cheapest VM with x CPU and Y memory" API
Depending on your workload you might be able to actually substitute a single 8xlarge with two 4xlarge for example... A while back I was actually doing something like that to save some money :-)
They do have the same amount of memory (and twice the CPU). But if you run a workload that automatically scales to the number of available cores, starting twice the number of processes / threads might well run you out of memory.
The article is interesting, but blindly running your code on unexpected instance types may be more "exciting" than the author makes it sound.
We see individual spot instances go away every few days which works pretty well for GKE. The older preemptible class of instances restarted every 24 hours which was more of a pain (mitigated a bit with a preemptible killer to spread the restarts out).
Some workloads might not. You wouldn't want to run stateful workloads on spot, for instance. In our case, we have something that doesn't handle bootup under load very well, and until we can improve that, the overall reliability is not as good.
I also like GCP's way of pricing these: you say whether your workload is preemptible or not, and you get discounts. You automatically get discounts if you run the workload for a long time.
There are differences, especially if you actually want to take advantage of the additional resources. Once I learned how to use the d variants, to mount the NVME disks for maximum advantage, to deal with the lack of persistence, and so on... the d spot instances were a steal! A lot of up front cost to use them properly, but performance was excellent and the price remained at the market floor almost 100% of the time. In the long run it was cheaper than the same machine without the fast disk attached! I assumed someone would figure it out and my advantage would disappear as the market corrected.
But I underestimated the sheer number number of ec2 instance types to evaluate! While it's true, having a diversity of instance types will help you keep spot costs down and ensure you're never priced out of that one blessed instance type - There's an upfront cost to make that system work though. Engineering teams just don't have the time to test on them all. At worst there are hidden differences that negatively impact performance, at best you might fail to take advantage of the resources and leave them idle (can your app even push 25Gbps or max out an NVME RAID).
We're heavy users of spot here in Intercom. I spot-checked our biggest workload, and this week we could have paid around 10% less if we were able to get the cheapest spot host possible in us-east-1 that is suitable for our workload (all 16xlarge Gravitons). However that would be at the cost of fleet stability, I think that to run relatively large production services used in realtime on spot you need to prioritise fleet stability, so choosing the "Capacity Optimized" strategy. We've seen incessant fleet churn when trying out cost optimised strategies.
I found it easy enough to do that in one region, but I've got some compute workloads that just read/write from S3 and are not latency sensitive.
They do need 128 GB RAM and ephemeral disks.
> need 128 GB RAM
Eh?
Not that they didn't do anything memory intensive.
I just want to specify the allowed instances in my CreateInstance command. Create an instance in these subnets, with these allowed instances, preferably a spot instance, but if none exists I’ll be happy with a normal one.
Yeah that's probably what's going on here. Complexity & that its just a bit counterintuitive
It’ll look at your pods’ cpu and memory requests and choose an appropriate instance type for you, and the cheapest spots where appropriate
For example, let's say you need to process a few million images, each taking a few seconds to process. You can start a manager task that distributes images, and a pool of interruptible workers, when a worker dies you just reissue the images to another.