Farewell to the Era of Cheap EC2 Spot Instances
pauley.me
pauley.me
I gave up on ec2 when they started requiring you "request for quota" to start a gpu instance.
You have to "request for quota" even if you want to run a single instance.
You have to specify which specific instance type you want.
You have to specify which region you want it in.
Then you sit back and wait for some human to "approve your request".
In my case this took 24 hours.
In this time I literally could have walked (not driven) down to the local computer store, bought a computer, pushed it back to my house in a shopping cart, spent a few hours configuring it and still be left with 16 hours to have a sleep, eat and do some other things.
AWS quota system is so far from "scalable" and "elastic" that it's effectively useless. You can't design any sort of infrastructure around that sort of quota system.
I dumped AWS at that point. Mind you, Azure has exactly the same quota system.
Seriously, just rent a server from Ionos or Hetzner or get a fast Internet connection and self host. It's faster and better and cheaper than any of the big clouds.
It's not impossible but quite technically challenging to find out how much you are actually paying for a spot instance.
It felt really dodgy to me that I might have started a spot instance thinking I was paying the minimum listed rate but somewhere amongst the formulas and terms and conditions AWS actually decided I would be paying maximum rate.
They don't actually tell you up front what your spot instance are costing, no doubt it would not be hard for them to do so, but it seems a deliberate strategy to hide this information.
It's just not worth the game playing when self hosting or renting a server from Ionos is cheaper and takes away the uncertainty.
Which is?
I don't really buy the idea that companies can't afford to run their own computers due to hardware and infrastructure costs, and that clouds don't require hiring or specialised technical experts.
There’s much less value if you never need to scale of course, which is true of 90% of businesses. But doing a dumb experiment on a 10 cent spot instance for an hour where it would’ve taken a week to arrange space on the company VMWare box? That is awesome.
It's only worth it if the pay-per-second box has an equivalent or better cost-performance ratio than the "legacy" alternative, otherwise it's still cheaper to get one "legacy" VPS/dedicated server, use it for 10 seconds and let the rest go to "waste".
Sadly, the scourge of crypto mining combined with credit card fraud makes it really hard to do otherwise for any sizeable hoster. :/
For example, a theoretical "prepaid AWS" might allow you to put a hold on a vCPU-month of account credits to start a 1vCPU instance. But what about the bandwidth egress fees when someone makes requests to said instance? Those are going to be completely variable, depending on how much traffic the instance receives.
As a nobody, I want to use AWS, but I refuse to have the unlimited liability in case I screw up something and wake up to a $30k bill. Hell, they could even over-charge me on the credits, and I would gladly take that deal if I knew that once the kitty ran dry, services would stop.
Would there be edge cases and complications to resolve (eg what about storage?). Sure, but AWS pays some smart people a lot of money to figure out tricky things.
The answer for you is LightSail.
LightSail is a standard VPS. But if you want to upgrade to “real AWS” later on, you can. The only thing that I’m aware of that could cause bills to go up is egress over your allowance.
To be honest with you, I would be slightly afraid of screwing up and having an unexpected bill from AWS if I were doing a personal project and I do this for a living. There have been plenty of times where I left something expensive running or provisioned an expensive service (Kendra) and forgot to shut it down until I ended up on a list of “people with the highest spend” on our internal system on one of my non production accounts.
But AWS prefers profit.
AWS has over 200 services. How would you implement that conceptually?
I know it’s a real concern when learning AWS for most people. I first learned AWS technologies at a 60 person company where I had admin access from day one to the AWS account and then went to AWS where I can open as many accounts as I want for learning. So I haven’t had to deal with that issue.
But what better way would you suggest than LightSail where you have known costs up front?
1. spending limits are fine-grained — rather than having one global budget for your entire AWS project, instead, each billable SKU inside a project would have its own separate configurable spending limit. The goal here isn't to say "I ran out of money; stop trying to charge me more money"; it's rather to say "I have budgeted X for the base spend for the static resources, which I will continue paying; but I have budgeted Y for the unpredictable/variable spend, and have exceeded that limit, so stop allowing anything to happen that will generate unpredictable/variable spend."
This way, you can continue to pay for e.g. S3 storage, while capping spend on S3 download (which would presumably make reading from buckets in the project impossible while this is in effect); or you can continue paying for your EC2 instances, while capping egress fees on them (which would presumably make you unable to make requests to the instances, but they'd still be running, so you wouldn't lose the state for any ephemeral instances.)
2. AWS "eats" the credit-spend events of a billing SKU between the time it detects budget-overlimit of that billing SKU, and the time it finishes applying policy to the resource that will stop it from generating any more credit-spend events on that billing SKU. (This is why this kind of protection logic can never be implemented the way people want by a third party: a third party can only watch AWS audit events and react by sending API requests; it has no authority to retroactively say "and anything that happens in between the two, disregard that at billing time, since that spend was our fault for not reacting faster.")
Note that implementing #2 actually makes implementing #1 much easier. To implement #1 alone, you'd have to have each service have some internal accounting-quota system that predicts how much spend "would be" happening in the billing layer, and can respond to that by disabling features in (soft) realtime for specific users in response to those users exceeding a credit quota configured in some other service. But if you add #2, then that accounting logic can be handled centrally and asynchronously in an accounting service which consumes periodic batched pushes of credit-spend-counter increments from other services. The accounting service could emit CQRS command "disable services generating billable SKU X for customer Y starting from timestamp Z" to a message queue, and the service itself could see it (and react by writing to an in-memory blackboard that endpoints A/B/C are disabled for user Y); but the invoicing service could also see it, and recompute the invoice for customer Y for the current month, with all spend events for billing SKU X after timestamp Z dropped from the invoice.
1. Take cash, load it onto prepaid cards.
2. Buy GPU instances.
3. Mine crypto.
4. Pay for compute with #1
5. Sell crypto
6. "Clean" cash
I just don't have the mind of a criminal I guess, it sounds like this was figured out a long time ago.
There are (were) enough random kids in thier basements making millions on crypto that it probably wouldn't turn too many eyebrows either by the IRS if you actually tried to report your crypto earnings cleanly, they would likely spend their efforts chasing down the people who weren't reporting.
On other hand credit card fraud is real problem especially due to fact that Amazon don't want to add any KYC burden on their customers or try to track down any anomaly spending. After all good chunk of AWS profits come out of fact how bad people of tracking their cloud costs.
If Amazon to implement built-in option to track suspicious jumps of AWS costs then it's not only gonna cut on fraud, but on overall AWS profits.
OK. https://aws.amazon.com/aws-cost-management/aws-cost-anomaly-...
the downside is the insanely high service fee -- 7 or 8 bucks per 500
Basically for money laundering you want to have a good story for your money that you can get tax authorities and banks to believe. "I found some crypto" would only work at small scale, and crypto is considered "high risk" anyway so you'll get higher scrutiny.
If I was in the business of laundering money and I wanted to do it via crypto, I think I would create an NFT instead then buy it from myself. Then I have funds I can transfer to cash, and I can claim I got them through my great artistic abilities.
When I use university-linked accounts, I get insta-approved for almost whatever I want. (2,048 spot T4 GPUs? Go for it.)
When using my personal CC-backed account, my max quota is often an order of magnitude lower.
They have a quota system for sending emails too, and it's not because they need to purchase any hw for sending more emails. It's because those are also a magnet for hackers.
My bank contacts me if there are any questionable transactions.
Surely AWS is capable of this?
Yes, (most of) the detail is in the CUR/CBR - but you have to be smart enough to understand it.
It's misleading to say that AWS will tell you all you want to know. Try getting them to explain your network costs in detail using network load balancers in a simple email.
I'm sure they're capable of many things that they don't actually want to do.
But if weird stuff happens, AWS will take action. I'm sure part of it is because they care, but another is because they don't want to get stuck holding the buck.
I don't agree. Their reality distortion field is fooling customers, AWS has their cake and eats it too. Try being a startup and spin up 1000 c6in.8xlarge for a 2 hour batch job.
I hope somewhere in those 16 hours you bother to return the shopping cart. Or are you one of those that just leaves it where ever because you can't be bothered? I will be bringing this up at the next tenant's meeting.
http://blog.tyrannyofthemouse.com/2016/02/some-of-my-geeky-t...
For a list of cloud GPU providers with rough price comparisons, I created this page: https://cloud-gpus.com
Given the prices of cloud infrastructure, it doesn't take much before walking to your local computer store and buying a GPU (or more!) becomes more cost-effective.
They closed last year, it's now a 50-minute bus ride and then ~10 minutes walking (one way). Apparently we were one of the last customers to buy most parts for a desktop there. I fear we may not be able to refer to corner stores much longer if they're not of a type most people actually need on a weekly basis, with things moving to online sales
> In my case this took 24 hours.
Even shipping is faster than that, so yeah point still taken
Well there's your problem. You don't need to sit back the whole time while you're waiting for approval, it's okay to leave your chair.
Jokes aside, I think this probably depends on use-case. You can't easily return the computer. And you can't easily buy (and then return) 100 computers if you have a large one-time computation to run.
Effectively, it starts up cheap spot instances (based on specified criteria) across a variety of instance types to replace whatever regular instance in an autoscaling group comes online and then spins down the regular instance.
EG: That m4a you wanted may be expensive... but nobody is using m4ad so it's 85% off and it meets the specified CPU/RAM requirements... auto spotting will spin it up instead.
Having used it on and off over the years it is sometimes eyebrow raising to see 4xl boxes running cheaper than the xl box they replaced :)
Are you sharing this as an experience of using AutoSpotting or other mechanisms to launch Spot instances?
What you described may still happen occasionally and it happened a lot with older versions, but the recent versions of AutoSpotting, especially the commercial edition available on the AWS Marketplace (although the community edition available on GitHub also does it to some degree) should avoid such situations in the default configuration should be much more reliable than before.
Actually when it comes to capacity, when Spot capacity is not available and we failover to on-demand, the latest AutoSpotting version will failover between multiple instance types with the same diversification used for Spot instances so you're pretty much guaranteed to have capacity even better than your initial single-instance ASG.
We also use the capacity-optimized-prioritized allocation strategy with a custom priority based on lowest cost, but also preferring recent instance types if available for more performance and lower carbon emissions.
Let me know if you have further questions about AutoSpotting, Im happy to help.
> In reality, instance preemption is rare in most instance families
I have often seen the following:
1. The demand for spot instances increases for a period of time.
2. Supply is inelastic so prices quickly rise to the full bid amounts.
3. Most users are unwilling to go hours without any spot instances running, so they use very high bids.
You either have to tolerate hours of "down time" every few months, or pay exorbitant prices during once in a while.
On multiple occasions, I've seen the spot instances priced significantly higher than on-demand.
EDIT: Apparently AWS changed the Spot pricing a few years ago though. My anecdotes are 5 years old.
How is the cap price not on-demand cost? Why would you not just swap to on-demand at the turning point?
This means "bidding" more won't protect your capacity, AWS will just take it if they need it.
Prices are now relatively flat, and changing slightly over time by a fraction of a cent per day with a cap at the OnDemand price, while previously there used to be lots of from 0.1x to 10x the on demand price and with swings even multiple times within a day.
It is because many systems are only built to support spot. When running a workload, they might not want to give up availability so they just pay the higher price. To be fair, this is a game that AWS engineers. They essentially want to kick people off the infrastructure to free it up for other uses. So this is a way to signal how badly you don't want to be kicked off. A lot of times after a price goes above on-demand pricing, it will dip down below fairly quickly as workloads all re-adjust. Some companies are willing to play that gamble, that a short period above on-demand is still cheaper in the long run when you average out the invoice at the end of the month.
Most hours, you’d pay something close to $0.30/hr. Some hours, you’d pay over a $1.00/hr, but you’d save money overall against on-demand.
(This ignores Reserved Instances and Savings Plans.)
If you set up an EKS/K8s cluster and install Karpenter. You can configure it very easily to use spot instances anytime prices are available for less than on-demand and to use on-demand when spot instances are unavailable or too expensive.
You end up never thinking about it but having full availability.
In practice it means the ceiling on your bill is the on-demand price, but you usually average out much lower.
I suspect as more and more people switch over to this model then the use of spot instances will stay closer to full saturation, with those discounts becoming negligible.
They have a whole section in their docs about this: https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/spot-int...
It's all about how your spot instances can be terminated. Perhaps this could interfere with a workload that might take a long while to finish.
On paper your app should be resilient to these things but in practice it's essentially an unforgiving chaos monkey that will delete servers from your cluster.
If you can handle that, you can definitely save a lot though. Especially web servers where it's expected you'll be responding in a few seconds while maybe your background workers are churning through longer lived tasks. The problem there is you need to already be big enough to have different nodes handling web vs worker traffic. A lot of decently sized apps can operate in a world where maybe you have 2-5 nodes on your cluster and your web and worker pods land on any of them based on Kubernetes' default scheduler.
This does require architecting your services to be shut down tolerant, but that's par for the course. If you're using Kubernetes you've probably already settled on the "[herd, not] cattle, not pets" idea.
If you have heavy duty work being done, it might get kill -9'd after the grace period is surpassed. If you weren't using spot instances you have full control over the grace period.
You can of course handle this too, by breaking up your work in ways where it can be interrupted without losing progress but depending on what you're doing this could be very complicated. Basically the takeaway is you can't blindly start using spot instances, even if your app has been running in Kubernetes successfully for a while.
One possible cause of this trend is actually still consistent with spot demand growth in aggregate: when customers deploy spot fleets with on-demand backup capacity and all spot prices get closer to on-demand, those fleets would become more likely to revert to on-demand, rather than a more expensive pool of spot instances.
Measuring this would be a great opportunity to leverage the vantage point that your team has!
You simply run Spot from other instance types, and if no Spot capacity is available you run at reduced capacity.
Alternative solutions such as AutoSpotting.io are doing it and allow you to run with failover capacity.
Otherwise they would be no different from any other instance type.
I have two home workstations that are collecting dust. I have 6 and 16 core CPUs with HT hooked up to 32 and 128gb ram.
I figured out how to make my dynamic IP still serve requests accounting for IP changes using a background script. If someone has a better solution I’d appreciate that.
Both these machines are more than enough to do everything I could dream of building and they host dozens of my little apps and services that I use for myself.
I will only use cloud when necessary for commercial product and services.
- Rent a cheap VPS with an static ip address (for example, the most basic at Vultr)
- Set up wireguard or tailscale/headscale (or your preferred VPN) both on the VPS and your workstations, so you can access the services in your workstations from the VPS through the VPN
- In the VPS configure Nginx/traefik/caddy (or your preferred reverse proxy) pointing to the services in your workstations
Ta-da! Now people access through your VPS, but the processing is done at home. The VPS is cheap, and it doesn't matter if you home IP address changes. You could even take your workstation, move it to any place with internet connection, and it would keep serving pages without changing anything.I'll try setting that up over the weekend. Just wondering - would you recommend using your suggested setup for a "production" app?
Considering the hardware I own is powerful enough to host postgres + a half dozen services to run the app including a bunch of real time processing. If my calculations are correct I should have enough resources to handle a few thousand or more concurrent users, and at an old gig we were able to service millions of DAU with even less power. As long as I can meet <500ms client request times I think it should be fine, considering there's gonna be extra network hops.
Home broadband tends to be very asymmetric, where you may have 200-500mb down(into the house), but only 30-40mb up(out of the house).
1200 mbps down and 1100 up.
You will get a lot of people who will try to find all the theoretical reasons why this wouldn't work.
My recommendation is that you can try it out and use it until you encounter those problems, and then you reevaluate the situation.
It's not a theoretical.
Depending on where you live either:
- The available tech (i.e. without fibre to the premise your "last mile" might be entirely copper, or just the run to the cabinet, but both can affect speed) - ISP offerings (Some only offer symmetric speed on much more expensive "business class" packages)
can limit the upload speed you have.
I know this started as an anti-cloud stance, but there is a lot that cloud servers can do for the DIY self-hoster, for really not that much money per month. And one of those things is knowing that even though you're having a power outage at home, that your server is still up.
cloudflare costs $0, scales up to your gigabit connection without blinking, and there are docker containers that keep the IP updated like it if were dynamic dns service.
You'll have to explain what you mean by "intentional" given raising prices above market equilibrium would leave compute idle (again, very unlikely). In that way amazon is just a market actor (albeit a large one) just like customers.
Had a couple of calls with their “please don’t leave us” team and they cut our bill more or less in half (more through fixing my boneheaded designs though).
Why would that change?
But it was changed a while ago, so prices are now set algorithmically based on predicted demand/supply.
We run stateless calculations on EC2 across regions, and we definitely see that instances are harder to come by. Especially instances with GPUs. And for many instance types, the price advantage of EC2 Spot compared to committed spending is not significant anymore.
Basically, one of the problems was customers who just set the spot price to 10x of the nominal price and leave the bids unattended. This was usually fine, when the price was 0.2x of the nominal price. But sometimes EC2 instance capacity crunches happened, and these high bids actually started competing with each other. As a result, customers could easily get 100 _times_ higher bill than they expected.
Another issue was that EC2 spot internally in AWS was implemented as a "bolted on" service that was not on the EC2 instance launch path, so it was easy to game the market. You could do very fun things, like:
1. You need an instance type that is right now under heavy contention. Not to worry! There's a way to get these instances!
2. You create a small VPC with only a couple of available IPs.
3. Then you submit a thousand EC2 Spot bids at 10x the price.
4. Your bids win and EC2 terminates other customers' instances. After all, you're willing to pay more!
5. EC2 then tries to launch these 1000 EC2 instances into the VPC. And fails, because there aren't any IPs available (see item 2).
6. Whoopsie. The bids are cancelled and EC2 instances are returned to the pool. Oh, and you're not charged anything because instances failed to launch.
7. Profit! Now there is plenty of capacity and you can submit bids at a normal price.
Thanks for sharing, though I'm confused about this part. It seems like expected behaviour?
Any possible spot/auction system I can think of would have the inherent possibility of a sudden surge in pricing.
But it turns out that this is not a customer-friendly behavior. So AWS decided to remove the bidding out of the equation, and instead terminate instances based on a complicated scoring system.
The idea is that it's easier to deal with the missing compute capacity, which you notice right away, rather than be blindsided with a 100x bill at the end of the month.
Why is it happening to the U.S. regions especially?
It doesn't mention either that price-capacity-optimized allocation strategy.
If you want to run spots in production, please read about three two. There is no major drama if you apply them.
Most of the percentages are just wildly incorrect.