Cutting down AWS cost by $150k per year simply by shutting things off
tuananh.net
tuananh.net
I once found a “test” db cluster from an engineer who hadn’t worked in the company for 3 years. We were paying 300k yearly for it before discounts. It took me a literal click to shut it down. And I’m not proud of it but, had to send out an org wide email on the savings achieved (corporate politics :shrug:).
if you want to keep it up, you have to tag it.
once you tag it, it can opt-out 7 days. then you have to extend it (simply chat with our bot)
Create a rule and shame them on Slack.
Most places I’ve worked had no formal production readiness review before launching infrastructure.
>Show me the incentive and I will show you the outcome
Would print this phrase on every angle of the offices
Nobody in engineering cares about spending because there is no benefits on doing so.
Even more: most people are on fixed salary and will get paycheck anyway no matter how low their effort is.
>>Show me the incentive and I will show you the outcome
And in many cases this is a huge net win! After all, there's another way to waste company money invisibly: design a process which requires meetings and waiting while work is held up.
It basically helps keep things clean with decomposition and doesn't necessarily hamstring older devs as much while giving a good guideline for younger devs to work in. All things considered, it seems like not a bad system to me, and the team customizing the process to their own needs is nice as well.
There's a million ways for it to go wrong, but it's not too terrible on the whole I thinks. <3 :"))))
There is no such thing as "no process"; is something you always have whether you talk about it or not. The often heard "I hate process" is counterfactual then - what it really means is "I hate process that I see as intrusive/wasteful/whatever".
The +'ves you are listing are what comes from looking at how things are actually done and doing some of it a bit more thoughtfully.
The common -ves often come from stakeholders outside of developement injecting their needs ... sometimes this is unavoidable (e.g. regulatory) sometimes it is just political, but either way there are better and worse ways to do it.
(Escalating costs of pre-approval, and the need to design around every possible objection, are a big part of why physical infrastructure costs so much more in the West!)
An engineer either wears a striped hat and drives a train, or, went to a credentialed school and passed a bunch of test and is allowed to sign documents that state "this thing, if built this way, won't collapse and kill people."
It is expected that an engineer can predict with reasonable accuracy the expense and timeline of a project, and how to maintain the resulting thing, without resorting to voodoo like "scrum velocity." In large part that's because engineers stick to doing things that are well understood and predictable, and if there's risk they resolve the risk before undertaking the project. (Is there bedrock over here upon which to build a foundation? I don't know; let's find out first!). Sure, there are engineering disasters even today -- buildings that unexpectedly lean over and door/wall things that unexpectedly fly off the side of airplanes, but those are typically organizational / process problems not "engineering doesn't work" problems.
I’m calling BS on this. If this were true, we’d still be a ground species. Engineering has been and will always be about creating something electrical, mechanical, computerized, or all, that solves a problem. Understood or not. Engineers are not oracles. They can not predict whether a tower built in Italy will eventually begin to lean due to erosion. They can not predict that a steel beam rated for 300T of force would break at 180T. They can not predict a rogue developer removing a package from underneath their dependency tree.
You can give estimates all you want but you are still guessing.
If engineers were as you say they are, we would never have delays, we would never have traffic jams, we would never have crap software, we would never have flight.
"A cpu is literally a rock that we tricked into thinking."
But looking at the malarkey that goes on in "software engineering" or whatever -- clearly not engineering, at least not where I've seen it.
Engineering: a process of repeatably solving an understood problem predictably.
Craft: a process of solving an understood problem.
Science: a process of solving a problem without an exactly understood outcome.
Art: a process of working.
These are all made-up definitions.
I'd expect a software engineer to give me a system that locally caches and verifies distribution artifacts and validates changes -- a craftsperson who gives me a tool chain that yeets goo from the internet and builds on that without validation is not, in fact, an engineer. They could be quite practiced at the art of building working systems, but they're not managing risk....
We call it help desk, not engineering.
"Any idiot can build a bridge that stands, but it takes an engineer to build a bridge that barely stands."
Ahhh - that old craptacular definition. You completely ignore mechanical engineers, chemical engineers, electrical & electronics engineers. Not all engineers make bridges.
Secondly, the implied cause and effect even within civil engineering is a fantasy. Signatures on documents by credentialed engineers doesn't prevent disasters as you noted: Bridges fall down, buildings burn. Read the engineering reports on civil engineering disasters, and look at the consequences for the engineers involved.
You do some handwaving about organizational/process problems, but actually that is the key to safe engineering. Organisations deliver engineering projects and they do it across jurisdictional borders using insurance and liability and with a variety of other means that work: "signatures don't prevent disasters".
Lockheed Martin's skunk-works and SpaceX are real engineering. Any good definition of engineering needs to encompass an extremely wide variety of activities.
Engineering is compromise. I have no love for Musk but him saying build that actuator for less than $5k is actually true engineering: https://news.ycombinator.com/item?id=39085892
I would like to know the psychology behind why people wish to believe credentialed signatures are so powerful? Maybe a cross between two concepts #1: "that individual engineers run the world" and #2: "that retributive punishment of individuals works as a deterrent". I think concept #1 comes from the egotist idea of most engineer-types that we are the center of everything (I need a whole article to explain the concept). I think concept #2 is related to beliefs about the value of incarceration and also punishment beliefs derived from religion (especially in the USA where prisons are not fixing problems?).
Edit: issue #3: the idea that we should make rules about what words mean. It takes a certain worldview to think words should be defined rather than evolve (or worse that words should be part of a justice system)
I suppose you've got an engineering degree in pedantic engineering? Engineers manage cost and risk. The skunk-works stuff is marginally "science" not "engineering" given the relatively large budgets and relative lack of "we know this works." Cern is similarly an enormous engineering enterprise in that it's a huge stack of "we know this works" in service of "we're not sure what this will do"
A discussion of how "software engineers" deliver projects with neither cost or risk as part of the process implies, to me, that they're not engineers.
I provided counter-examples that show engineering encompasses a lot more than your definition.
I simply don't understand why anyone thinks writing software is somehow uniquely not "real" engineering. Somehow we are indoctrinated to believe that it isn't but all the evidence seems to show software engineering is a valid description.
I have no lack of experience watching the fuck-ups made by electronics engineers, or the fuck-ups made by mechanical engineers. You appear to want to define engineering only as certified civil engineering. And I've seen enough of their fuck-ups too, with signatures. In fact I'll ask my bridge engineer friend from uni about it! Unfortunately my bridge building grandad is dead so I can't ask him.
A site reliability "engineer" or a software "engineer" is not an actual engineer just because they've got that in their title or job description. If I were to hire a "chemical engineer" position and instead hired a chemist, or a mechanic, or a rando who's cooked meth, I may end up with things working okay, or I might end up with a serious mess, even if those people I hire call themselves "engineers" (but in fact have no formal training as such).
I'm not sure to what degree credentials matter, but do credentials matter more than "not a god damned bit" ?
I'm not saying the title makes you "not an idiot" -- people gonna people -- but attention to "cost" and "risk" is (theoretically) one of the distinguishing characteristics of engineering training vs ... "mather" or "programmer" or "philosopher".
I have a bachelor of engineering title I can use with my name, but that is another distinct type of bullshit.
In New Zealand one relevant legal certification is CPEng which you can apply for after receiving your degree and working for a few years: https://www.engineeringnz.org/join-us/cpeng/ And apparently our government agreed in 2022 to introduce a new licensing regime for engineers doing safety-critical work.
But in an international world, how relevant are certified individuals? When I purchase a stove from a US brand and it catches fire, there needs to be other liability/retribution/corrective systems to deal with the problem. It matters little to me who signed off on the product in the US.
Can I import custom structural steel beams? How many New Zealanders have signed off on this steel construction: https://ccc.govt.nz/the-council/future-projects/major-facili... We need a new stadium because the last one broke. Unfortunately it wasn't insured due to some cockup at the city council (which I suspect had zero retribution on the people that cocked up - I wonder if they signed bits of paper?).
Over-credentialisation is a problem too - where is the right balance? The shift to everyone needing credentials is fucked. My friends (nurses, teachers) literally weep at the absolute trash they have to "learn" for their credential. I also vividly remember the crap I needed to disgorge to get my degree.
I don't know what the answer is, but I honestly believe most credentials are pointless waste and adding more credentials is not actually effective. Neither do I believe that that the anarchy of libertarian free markets are a workable answer.
1. Don't ask developers how much something costs, engineers love optimisation, getting as much as possible out of a system for cheap is great fun.
2. Lock down the UI, so devs can't even find out how much things cost. That's my current situation. Why block the billing dashboard, then expose it through billing dashboard tools that are not really any better, and in many ways worse?
It's rhetorical really as I know why. Terrible architecture from "enterprise". Stick everything in a single account so it's hard to figure out how much is your spend. All 3000 databases, and make sure your k8s cluster is 5 8XL boxes so no one can scale down excess capacity.
Classic I Burn Money consultancy!
This is so true. billing transparency is very important.
in the past, i had a case like this: dev accidentally enable backup policy for test database with no retention. finops think that db backup is important and ignore it. dev has no access to billing and have no idea what's creeping up the bill
> You need to request access every 60 days.
Luckily it's now added to my permanent role, but even then no billing access? FFS.
It's easy not to grant engineers access to the billing dashboard.
It's easy to put everything in the same aws account.
Inside Amazon, we're supposed to set up new aws accounts for every service and realm, so we know how much X service's beta environment is costing
Plenty of ways of doing that, like making a cross account shared VPC for example.
Everything is still accountable.
In certain circumstances, absolutely, however it's extremely aggravating to be in the position of being constantly pestered to ship features faster without the authority to overprovision some of the infrastructure the software runs on.
Waking up in the middle of the night because we saved money by allocating too little disk for the primary database or because the latest release included new dependencies that increased memory usage and the OOMKiller is picking off web servers like a wolf in the lamb's pen, or we're just swapping our way to hell while web requests 502...eh. Not for me.
More visibility into costs, though, absolutely agreed. Engineers should know that when they turn on some new cool serverless gizmo and then forget about it, it's costing $ each month.
I didn't mean that engineers love having no control over their systems, I just see the labours of love that get posted here about getting nginx to throw out 1000 pages a second on an Atari 800, or getting LLMs designed for $2000 GPUs running on a phone.
The question should be, we currently cost $X a month and we need to half it because [reason], what can we do to bring it down? Which might be reducing hardware, or maybe something else, might be both. Puzzles can be fun.
That’s because they were explicitly told not to worry about costs for the last 10 years so majority of ICs at this point never had to do it their entire careers
And we celebrate costs slashing as much feature delivery and other stuff.
But this is entirely a management problem: at my previous job, only one manager (skip-level manager from my point of view) knew what exactly were we paying for infrastructure.
That moron wouldn't share that information with us engineers managing infrastructure of course, so there were a lot of infrastructure choices that didn't really made sense according to the public prices but (I guess?) made sense according to a price sheet we didn't know.
So we didn't know what we were spending, didn't have the basic data to estimate the price of a new solution or a new service and didn't have the data to determine how much would we be saving by making changes (optimizing stuff etc).
I fought that battle for a bit but then i just said "GFYS, i'm not going to have fights with you so that you can save money" and let go. Later i left the company completely.
Former colleagues tell me it's even worse now: there are consultants from the cloud provider involved, they know the pricing deals, and whenever the topic comes up the manager shushes the consultant so that the engineers don't hear the prices.
tl;dr: it's an entirely artificial problem, and it's most likely a cultural/management problem.
edit: and i'm not even talking about incentives, as somebody else has correctly pointed out.
And that's precisely why you and your little bootstrapper or indie firm should not be using globocloud: you do not have mountains of cash to piss away. Bare metal is trending again. And in this downturn, it's no wonder why. Smaller companies are getting smarter and more efficient. They've decided to chase money instead of cargo cults.
Globocorps burning cash on globocloud is not a signal for small fish to do the same - it's a signal to do the polar opposite. You're not going to become like them by copying what they're doing now. It will not work for you. Globocloud isn't successful because they shovel cash into AWS's shredder, they shovel cash into the shredder because they're successful.
It's the same issue in any large organization: Large levels of success somewhere allow for large levels of waste somewhere else, but often the waste is not required for the success to exist: The success just makes the organization complacent.
I was laid off this week in a mass layoff because the company doesn’t have enough money to pay all of us anymore. It’s disappointing to see, and I wonder how many other teams ignored these optimizations and how much unnecessary total cost it all summed to.
One of the downsides with this approach is that engineers/developers are not very good business people and don't really understand the notion of "the cost of doing business". And from time to time we have issues with "but it costs $70 more per month", and spend $1000 to optimize those $70 :)
In the end, even with some of the wrinkles mention above it helps and saves money when costs are transparent and readily available for anyone.
I noticed that one of our S3 buckets had high data transfer costs, a bucket that our app downloads HTML+JS assets from when we push out a new release. I downloaded the "directory" of files for our latest release and saw it was mostly node_modules. I checked the code and confirmed that, yes, if this file exists in the bucket then it'll be downloaded by the user. I wrote a quick Python script to list out each directory that had this problem, and a quick Slack message to the appropriate team later, we discovered the specific commit that was the cause, a change to our CI that inadvertently uploaded that directory when we wanted to ignore it.
A few months later, I checked the billing metrics, the effect was an avg of $12,500 reduction in cost for this bucket, or around $150k per year, or 4% of our bill. Not bad for one hour of work. Over the course of a quarter I reduced our bill by over $1m, or around 30% of our bill.
I might write a blog post explaining how to go about something like that. A lot of people are not familiar with tools like Trusted Advisor which can easily tell you if you have, for example, unused EC2 instances that can be terminated.
I would be happy to receive some extra cash, don't get me wrong, but I work for non-monetary benefits as well, and I have received some of those as part of this work. If I worked at a company with a different culture and I was being punished for doing the work, I would demand some bonus.
A better way to have aligned incentives for the company and the employees would be to allocate a bonus pool for the entire company, from which AWS expenses are taken out of, but that might be a bit unorthodox.
Also a perverse incentive.
If we use ec2.small, the customer's query will take 3x longer but be half the price. Let's turn off the nightly security audits. We can live with quarterly backups, right? What do we need all these logs for, anyway? We could hack something that works together in 2 weeks, but if we spend 3 months, it could be really efficient, let's do that...
The incentives are designed to form an enormous cash siphon. From aggressively marketing toward fearful & liable (or maybe just tech-cost-illiterate) upper-management to the silencing effect that the low-rung experts experience when sounding the alarm.
Most time this "opportunity cost" is then spent on useless hacky features that are never used and forgotten right after release (redesign anyone?).
Of course there are exceptions to all of these, but IME the majority of companies are either focusing on pennies or ignoring it completely. Not sure why you don't see balanced approaches more often. Maybe this will change with less VC money flying around.
> (2) I can do this on my own time and keep some % of the savings for myself as a reward
This is a textbook case of perverse incentives.
Theoretically, your company could reduce their commitment to 4M next year, but the AWS sales would start negotiating hard against it, like "you will not get the same discount with less commitment".
Ironically, I was asked by a manager some time ago if we can imagine using some (more) resources from AWS to reach the next spending commitment. If you're just below the threshold, it's probably inconvenient.
EDIT: To give some context: you can only do this, of course, when you know that you're not really wasting resources, otherwise you end up with just burning money to save little to no money :)
Even if you have one year reservations on your instances, starting service migrations/deprecations now would pay off quickly enough. Your commitment expires in 6 months, on an average basis.
If they are paying you to saves 10s to maybe hundreds the company is losing money on you so they won't do this.
If your at a public company, look at your company's quarterly reports and see what it would take to many any kind of impact on net income.
I reached out to the team and they turned it off, it saved us $1m a year. The higher-ups rewarded me by telling me that a team should have caught this so I should meet with them now.
that's why we did finops dashboard first thing when we first started the cloud journey.
can't optimize if you dont know.
(A) "Hey wow, this is great! We are so excited to be saving from here on out." OR, (B) "This should have been caught earlier. $TEAM was supposed to be experts..." and then blame game starts.
It is really unfortunate when institutions react in the latter way. Often the engineers are assigned to cost optimization, along with a million other things. And, the incentives aren't really aligned well to reward savings. For example, S3 Intelligent Tiering is the right thing in 99.9% of cases - so it should be your default bucket type. BUT, engineers often face only downside risk for the change, and very little upside reward. And, it isn't their money so they just leave it. The cost of overprovisioned S3 can be staggering!
What is really needed is to establish a proper FinOps discipline, put someone in charge of cost savings, and make sure incentives are aligned properly. And of course check out CloudFix if you can!
The biggest problem they have is they have no business insight into what these costs are, and if we can reduce the cost without any kind of engineering, effort, or loss of performance.
I have zero insight into the costs.
Yes, my company could turn that on for me but it's rare that they do so it's nearly impossible to know if I did something that costs a lot of money (relatively or in general) without access to the cost explorer/billing dashboard.
And before "well can look up what a t2.2xlarge costs and calculate it", sure. In a very contrived example I might be able to see what it costs but so many things are hidden/hard to see in AWS. For example, I recently spun up an RDS customer on my own AWS account. After testing for a while I decided it wasn't what I wanted and I deleted the cluster. Fast forward a month and my bill is well over what I expected (Like $30, no it's not a ton of money but it's my personal account and I wasn't expecting that charge). Come to find out it created a VPC as part of the RDS cluster (I think maybe it was for the RDS proxy? Still not sure) that didn't get deleted. I had to go chase that down and even that process wasn't easy. I had to make sure that it wasn't be used by anything else and then delete other things that were created when I made the RDS cluster before I could remove the VPC.
I was only able to do the above because I had access to the billing info. I would have left that VPC indefinitely on my work's AWS account by accident and been none the wiser.
I'm more than happy to take costs into account but without access to what things are actually costing us I can't help that much. Mostly because I need to know the costs to know what's worth optimizing. Sure I know I could improve X feature but if that costs us pennies a day (or month sometimes) then it's not worth it. Similarly if I know feature/infra Y is costing $XX,000/mo then I know I should rethink or investigate if that's correct/worth it.
in the past, i had a case like this: dev accidentally enable backup policy for test database with no retention. finops think that db backup is important and ignore it. dev has no access to billing and have no idea what's creeping up the bill
I've asked, off-hand, a couple times for billing access but nothing has come of it. I don't want to seem pushy but also it feels like data I need to perform my job to the best of my ability (especially at a small company). I don't think it comes from a place of "We don't want to give Josh access" or secrecy as much as it not being a priority but I need to bring it up again.
I don’t believe the VPC was factored in when I used that calculator, even after selecting RDS Proxy.
1 - https://github.com/turbot/steampipe 2 - https://github.com/turbot/steampipe-mod-aws-thrifty 3 - https://github.com/turbot/flowpipe
The fact that they even outsource their compute to AWS is kind of surprising when they could just fill up their existing data centers (like VNTT https://vntt.com.vn/) with equipment, and save a whole lot more money.
My guess is that nobody in corporate approved this guys posting and if word got back, it would disappear quickly.
Reminds me to forward this to my buddies who run Timo, which VPBank used to own, but then dropped [1]. Timo was the first forward thinking bank in Vietnam with a great tech platform, likely because it was started and run by foreigners... ¯\_(ツ)_/¯.
[0] https://www.reddit.com/r/VietNam/comments/zvo553/sharks_ate_...
[1] https://fintechnews.sg/42738/vietnam/vietnams-challenger-tim...
This is the way.
A similar idea has been bouncing around in my mind for a while now. An ideal, turnkey system would do the following:
- Execute via Lambda (serverless).
- Support automated startup and shutdown of various AWS resources on a schedule influenced by specially formatted tags.
- Enable resources to be brought back up out of schedule when demand dictates.
- Operate as a TCP/HTTP proxy that can delay clients so that a given service can be started when it is dormant or, even better, the service isn't serverless but you want it to be. This can't work for everything, but perhaps enough things such that the need to run always on services is reduced.
Cloud Custodian [1] can purportedly do some of this, but I've been reluctant to learn yet another YAML-based DSL to use it.
So this is my "make things designed to be always-on serverless instead" project and the work AWS has done to make Java apps function on Lambda keeps me thinking about the potential to take things that 1) have a relatively long startup time and 2) are designed to be long running service loops, and find a way to force them into the serverless execution model.
My team mostly builds internal stuff and we save tons of $$$ by using Knative + Karpenter, which basically does that on container + EC2 levels.
If they say no then just go back to your regular duties.
I keep thinking I should be doing "cloud optimization" work and being compensated this way. Slicing and dicing output from usage/billing APIs and providing an "optimized spend" probably has the potential for a lot of low hanging fruit.
Ironically, I asked for a raise a year later and was denied, despite single handedly saving the company nearly $50000/yr. The raise I asked for wasn't close the cost savings I had brought. I left the company shortly after.
I saw someone else have a similar experience here and a comment to it was saying rewarding this produces a bad incentive...well, honestly why would I have even bothered cutting costs if I felt I wouldn't be rewarded? Not rewarding it just makes me half regret doing it at all.
- dependent services not coming up in the order you expect/want
- issues draining nodes due to crashlooping/erroring pods (can also be caused by dependent upstream services going down in wrong order)
- Persistent Volume retention/synchronization
- IAC not cooperating
- Configuration annoyances with deployments’ availability/replica settings
- Thundering herd types of problems
I can think of tons of things that can make this extraordinarily difficult. I’ve had many managers over the years pitch this idea of “rapidly deployable/destructable EKS clusters” and the projects always get killed due to the complexity around this. IMHO they simply aren’t really designed for this type of thing, however, I could be misunderstanding exactly what you’re trying to do.
This is exactly what we do: blue green eks cluster.
We just thought if we do it on monthly basis, DRP will be piece of cake :)
I’ve seen several clusters where one could kill more or less everything and it would just come back again.
Sorry for the rant, but this is usually wrong. The amount of people that just keeps their computer on is noticeable. And when I ask it's usually "just to avoid having to wait" or "I've always done that".
I personally always hibernate my computer. When I turn it off it takes more time, but I'm already on the other side of the building so I don't care. When I turn it on it takes basically the same amount of time, and it is exactly as I left it. People keep the computer on just because convenience...and I don't think it's a good thing.
- I have a plex server running on it
- I can remote into it from my phone, this comes in handy a lot of the time.
- I can remote into it when traveling through my Fire stick using parsec, which means I don't have to carry a laptop with me everywhere I go ( I also setup my phone so I can use it as keyboard/mouse when I do this).
Regarding energy costs, it's negligible for the benefits it gives me
Nothing keeps humming if it's not being used
90% sounds just about right. We are seeing figures going from $120/m for a VM-based QA environment to $10/m for a consumption-based / serverless stack.
We routinely see 10x savings when switching from RDS or Aurora. Especially if you start adding dev environments.
--> "The best optimization is simply not spinning things up!"
At least for local development and testing, as made possible by LocalStack (https://localstack.cloud), among other local testing solutions and emulators.
We've seen so many teams fall into the trap of "someone forgot to shut down dev resource X for a week and now we've racked up a $$$ bill on AWS".
What is everyone's strategy to avoid this kind of situation? Tools like `aws-nuke` (https://github.com/rebuy-de/aws-nuke) are awesome (!) to clean up unused resources, but frankly they should not be necessary in the first place...
Just set a cron to run the shutdown command with a grace period. And then if you're working late, you just cancel the shutdown and the shutdown will be retried in a couple of hours. And have a script or command to just run the cloud API calls to boot the VM in the morning / when needed, and the environment boots in a minute or two.
For other stuff I've been tempted to do a more complicated setup, with something like a micro-vm as a proxy, that will do the shutdown / activation on TCP connection, but haven't gotten around to it.
For people looking for how to save money on AWS - I'd [selfishly] recommend connecting up to Vantage. We profile AWS for all sorts of savings and give you the information on how much we can save prior to you paying us. It can be a good gut-check if nothing else on how well optimized you are.
Is there an offline method? I have not looked at vantage to see if it's possible.
Also, I've seen a lot of concern over blocked IPs, especially for lower-cost hosts. Is that an issue with Hetzner?
Still, Hetzner cloud is pretty good option, and there's more support coming on building on Hetzner.
https://www.ubicloud.com/ (from founders of citus) is mainly/currently targeting hetzner, for example.
Hetzner offers a 99.9% uptime guarantee only on their network. AWS has SLAs for every product offering - EC2 for example starts paying out credits if they fall below 99.99% uptime.
If you're a user of various managed cloud products, these will cost quite a bit to replicate on Hetzner and you'll be spending money on personnel to build these out and maintain them instead of just paying for the cloud product on AWS/GCP/Azure.
You need postgres? Use crunchydata postgres operator or cloudnativepg. Need multiple regions? setup wireguard.
IT's more work, but might not be a lot of work.
Cloud revenue in most large companies is at least 25$, maybe up to 40%, pure developer waste because nobody upstairs knows the difference.
Did I miss a new trend or something?
How do you reward cloud cost awareness without creating perverse incentives?