Not sure how to factor that $ into the equation.
2x FTEs to manage the AWS support tickets
3x FTE to understand the differences between the AWS bundled products and open source stuff which you can't get close enough to the config for so that you can actually use it as intended.
3x Security folk to work out how to manage the tangle of multiple accounts, networks, WAF and compliance overheads
3x FTEs to write HCL and YAML to support the cloud.
2x Solution architects to try and rebuild everything cloud native and get stuck in some technicality inside step functions for 9 months and achieve nothing.
1x extra manager to sit in meetings with AWS once a week and bitch about the crap support, the broken OSS bundled stuff and work out weird network issues.
1x cloud janitor to clean up all the dirt left around the cluster burning cash.
---
Footnote: Was this to free us or enslave us?
I assume whichever provides more margin to Jeff Bezos.
And no one cares about AWS certifications. They are proof of nothing and disregarded by anyone with a modicum of a clue.
I’m speaking as someone who once had nine active certifications and I believe I still have six active ones. I only got them as a guided learning path. I knew going in they were meaningless.
Having an AWS certification is not a requirement or even that important to get a job at AWS in the Professional Services department. Depending on your job position you are required to have certain certifications once you get there.
I now work for a partner and you are required to have a certain number of “certified individuals” to maintain partnership status. But even then, certifications never came up in my three interviews after getting “Amazoned” a couple of months ago.
But then again, after having AWS ProServe on my resume and having been a major contributor to a popular open source project in my niche, door opened for my automatically.
And based on my own personal interactions with other "certified" individuals, it doesn't actually mean anything.
What I don’t know is why AWS would rather pay a salesman $150k (or more… I looked up salaries a few months ago, but either way…) to sell the wrong things to customers, rather than have a software engineer who has actually used these products, sell the right thing to customers. I should hope that all AWS Solutions Architects need to pass the cloud fundamentals exam before interacting with customers, but maybe not?
Deming is rolling in his grave.
And even they aren’t to be confused with “consultants”. SAs are free to customers and give general guidance and are not allowed to give the customers any code.
Consultants are full time employees at AWS who get paid by the customer to do hands on keyboard work. But even we couldn’t work in production environments. We did initial work and taught the customer how to maintain and enhance the work.
If you don’t know the cloud fundamentals, learning enough to pass a few multiple choice questions.
As an anecdote, I passed the first one - the Solution Architect Associate - before I ever opened the AWS console.
I’m aware the bar is low when it comes to the entry-level certs, and that’s why I’d hope AWS SAs (the free kind) have to pass one or two.
One that's not immediately obvious is to keep on staff experienced infra engineers that bring their expertise for designing future projects.
Another is the option to tackle project in ways that would be to costly if they were still on AWS (e.g. ML training, stuff with long and heavy CPU load).
The total revenue so far for one cpu is 100x18x12x7 = $150k If used as a spot instance it’s 144/month, so about 200k
A standard i9-14700k gen has 32 threads, but it can run 12 of these instances (max 192mem). This CPU will cost you $800. Memory is cheap, so for about 1-2k you’re all set, and have a machine that’s way faster and cheaper.
Basically, buy a bunch of NUCs and you’re saving yourself around $1500 per month per NUC. It pays itself back in 1 month
Cloud hosting is —insane—
Not even touching memory ballooning for mostly idle applications.
Lastly, don’t give me reliability s an argument. These were all ephemeral instances that have no storage, so you’ll have to pay for that slow non-nvme storage platform.
Perhaps there are some good reasons to not choose such a provider once you reach a certain scale, but they now have their own versions of a lot of different AWS services, and they're more than sufficient for my own relatively small scale.
Also what happens at hardware end-of-life?
Also what happens if they encounter an explosive growth or burst usage event?
And did their current staffing include enough headcount to maintain the physical machines or did they have to hire for that?
Etc etc. Cloud is not cheap but if you are honest about TCO then the savings likely are WAY less than they imply in the article.
Your math is incorrect. The savings are per year. The job gets done once.
> Also what happens at hardware end-of-life?
You buy more hardware. A drive should last a few years on average at least.
> Also what happens if they encounter an explosive growth or burst usage event?
Short term, clouds are always available to handle extra compute. It's not a bad idea to use a cloud load-balancing system anyway to handle spam or caching.
But also, you can buy hardware from amazon and get it the next day with Prime.
> And did their current staffing include enough headcount to maintain the physical machines or did they have to hire for that?
I'm sure any team capable of building complex software at scale is capable of running a few servers on prem. I'm sure there's more than a few programmers on most teams that have homelabs they muck around with.
> Etc etc.
I'd love to hear more arguments.
TFA states that they maintain their AWS account, and can spin up additional compute in ~10 minutes.
If you're using less than a dozen servers manual configuration is simpler. Depending on what you're doing that could mean serving a hundred million customers. Which is plenty for most business.
I've worked at companies with their own data centers and manual configuration. Every system was a pet.
What we're missing is tracking what time is spent managing hardware and firmware, (or even network config, if we're being generous) and how much time is being spent on OS config.
From personal experience (as a sysadmin before it was entirely unsexy as a term) the overwhelming majority of my ops work was done in userland on the machine, maybe something like 96-97% of my tasks were nothing to do with hardware at all.
Since I got rebranded as SRE, the tools and the pay sure did get a lot better in the time, but the job is largely similar and ultimately running in VMs does make deployment faster, but once deployed I find the maintenance burden to be the same (or perhaps a little more) as things seem to become deprecated or require changes from our cloud vendor a bit more often.
> In the context of AWS, the expenses associated with employing AWS administrators often exceed those of Linux on-premises server administrators. This represents an additional cost-saving benefit when shifting to bare metal. With today’s servers being both efficient and reliable, the need for “management” has significantly decreased.
I also never seen an eng org where substantial part of it didn’t do useless projects that never amount to anything
A team does not use AWS because it provides compute. AWS, even when using barebonea EC2 instances, actually means on-demand provisioning of computational resources with the help of infrastructure-as-code services. A random developer logs into his AWS console, clicks a few buttons, and he's already running a fully instrumented service with logging and metrics a click away. He can click another button and delete/shut down everything. He can click on a button again and deploy the same application in multiple continents with static files provided through a global CDN, deployed with a dedicated pipeline. He clicks on another button again and everything is shut down again.
How do you pull that off with "Linux on-premises server administrators"? You don't.
At most, you can get your Linux server administrators to manage their hardware with something like OpenStack, but they would be playing the role of the AWS engineers that your "AWS administrators" don't even know exist. However, anyone who works with AWS only works on the abstraction layers above that which a "Linux on premises administrator" works on.
This only works that way for very small spend orgs that haven’t implemented soc 2 or the like. If that’s what you’re doing then probably should stay away from datacenter, sure
No, not really. That's how basically all services deployed to AWS work once you get the relevant CloudFormation/CDK bits lined up. I've worked on applications designed with high-availability in mind, which included multi-region deployments, which I could deploy as sandboxed applications on personal AWS accounts in a matter of a couple of minutes.
What exactly are you doing horribly wrong to think that architecting services the right way is something that only "small spend orgs" would know how to do?
How is an army of "devops" implementing your CF/CDK stack any different from an army of (lower paid) sysadmins running proxmox/openstack/k8s/etc on your hw?
My comment is really not about AWS. It's about the apples-to-oranges comparison between the job of "Linux on-premises server administrator" and value-added of managing on-premises servers, and the role of "AWS administrator". Someone needs to be completely clueless to the realities of both job roles to assume they deliver the same value. They don't.
Someone with access to any of the cloud provider services on the market is able to whip out and scale up whole web applications with far more flexibility and speed than any conceivable on-premises setup managed with the same budget. This is not up for debate.
> How is an army of "devops" implementing your CF/CDK stack any different from an army of (lower paid) sysadmins running proxmox/openstack/k8s/etc on your hw?
Think about it for a second. With the exact same budget, how do you pull off a multi-region deployment with an on-premises setup managed by your on-premises linux admins? And even if your goal is providing a single deployment, how flexible are you to put up this scheme to test a prototype and afterwards shut down the service?
Bullshit. I've seen people spin wheels for months/years deploying their cloud native jank and you should read the article - it's not nearly the same budget.
> Think about it for a second. With the exact same budget, how do you pull off a multi-region deployment with an on-premises setup managed by your on-premises linux admins?
You do realize things like site interconnect exist right? And it likely will be cheaper than paying your cloud inter-region transfer fees. You're going to be testing multi-regional prototype? please
Look there's a very simple reason why folks have been chasing public clouds and it has nothing to do their marketing spiel of elastic compute, increased velocity, etc. That reason is simple - teams get control of their spend without having to ask anyone for permission (like the old-school infra team).
1) not as reliable as you think you are 2) probably wasting gobs of money somewhere
I’ve set up an “RnD” account where developers can go wild and click ops away. I also set up a separate “development” account where they can test thier IAC manually and then commit it and it gets tested through a CI/CD pipeline. Then after that it goes through the standard pull request/review process.
Not everything is warehouse scale. You can serve tens of millions of customers from a single machine.
You don't click to start and stop. You start with someone negotiating credits and reserved instance costs with AWS. Then you have to keep up with spending commitments. Sometimes clicking stop will cost you more than leaving shit running.
It gets to the point where $50k a month is indistinguishable from the noise floor of spending.
I worked on a web application that provided by a FANG-like global corporation that is a household name and used by millions of users every day, and which can and did made the news rounds if it experiences issues. It is a high-availability multi-region deployment spread about a dozen independent AWS accounts and managed around the clock by multiple teams.
Please tell me more how I "never actually ended up with a big AWS estate."
I love how people like you try to shoot down arguments with appeals to authority when you are this clueless about the topic and are this oblivious regarding everyone else's experience.
The parent you're replying to resonates with me. A lot of politics about how you spend and how you commit, it's almost as bad as the commitment terms for bare-metal providers (3,6,12,24month commits). Except the base-load is more expensive.
It depends a lot on your load, but for my workloads (which, is fragile dumb but very vertical compute with a wide geographic dispersion), the cost is so high that a few dozen thousand has been missed numerous times, despite having in-house "fin-ops" folks casting their gaze upon our spend.
That doesn't preclude continuing to use AWS and other cloud service as a click-ops driven platform for experimentation, and requiring that anything that is targeting production to refactored to run in the bare-metal environment. At least two shops I worked at previously have used that as a recurring model (one focusing on AWS, the other on GCP) for stuff that was in prototyping or development.
That's part of the apples-and-oranges problem I mentioned.
It's perfectly fine if a company decides to save up massive amounts of cash by running stable core services on-premises instead of paying small fortunes to a cloud provider for the equivalent service.
Except that that's not the value proposition of a cloud provider.
A team managing on premises hardware barely covers a fraction of the value or flexibility provided by a cloud service. That team of Linux sysadmins does not nor will it ever provide the level of flexibility nor cover the range of services that a single person with access to a AWS/GCP/Azure account provides. It's like claiming that buying your own screwdriver is far better than renting a whole workshop. Sure, you have a point if all you plan on doing is tightening that screw. Except you don't pay for a workshop to tighten up screws, and instead you use it to iterate over designs for your screws before you even know how much load it's expected to take.
If you _actually need_ Kafka, for example – not just any messaging system – then your scale is such that you better know how to monitor it, tune it, and fix it when it breaks. If you can do that, then what's the difference from running it yourself? Build images with Packer, manage configs with Ansible or Puppet.
Cloud lets you iterate a lot faster because you don't have to know how any of this stuff works, but that ends up biting you once you do need to know.
Well said! At $LASTJOB, new management/leadership had blinders on [0][1] and were surrounded by sycophants & "sales engineers". They didn't listen to the staff that actually held the technical/empirical expertise, and still decided to go all in on cloud. Promises were made and not delivered, lots of downtime that affected _all areas of the organization_ [2] which could have been avoided (even post migration), etc. Long story short, money & time were wasted on cloud endeavors for $STACKS that didn't need to be in the cloud to start, and weren't designed to be cloud-based. The best part is that none of the management/leadership/sycophants/"sales engineers" had any shame at all for the decisions that were made.
Don't get me wrong, cloud does serve a purpose and serves that purpose well. But, a lot of people willfully ignore the simple fact that cloud providers are still staffed with on-prem infrastructure run by teams of staff/administrators/engineers.
[0] Indoctrinated by buzz words [1] We need to compete at "global scale" [2] Higher education
Anyone who says that hasn’t done it at scale.
“Infrastructure has weight”. Dependencies always creep in and any large scale migration involves regression testing, security, dealing with the PMO, compliance, dealing with outside vendors who may have white listed certain IP addresses, training, vendor negotiations, data migrations etc.
And then even though you use MySQL for instance, someone somewhere decided to do a “load data into S3” AWS MySQL extension and now they are going to have to write an ETL job. Someone else decided to store and serve static web assets to S3.
I specifically said "if it is designed well", and that phrase does alot of heavy lifting in that sentence. It's not easy, and you don't always put your A-team on a project when the B or C team can get the job done.
The article outlines a case where a business saw a solid justification for moving to bare metal, and saved approximately 1-3 SDE (depending on market) salary in doing so.
That amount of money can be hugely meaningful in a bootstrapped business (for example, for one of the businesses my partner owns, saving that much money over COVID shut-downs meant keeping the business afloat rather than shuttering the business permanently).
Source: former AWS Professional Services employee . I just “left” two months ago. I now work for a smaller shop. I mostly specialize in “application modernization”. But I have been involved in hairy migration projects.
> Our choice was to run a Microk8s cluster in a colocation facility
they go on to describe they use helm as well. there's no reason to assume that "a a fully instrumented service with logging and metrics" still isnt a click and keypress away.
your points dont make a whole lot of sense in the context of what they actually migrated too.
In a dream. In the real world of medium-to-large enterprise, a developer opens a ticket or uses some custom-built tool to bootstrap a new service, after writing a design doc and maybe going through a security review. They wait for the necessary approvals while they prepare the internal observability tools, and find out that there is an ongoing migration and their stack is not fully supported yet. In the meantime, he needs permissions to edit the Terraform files to update routing rules and actually send traffic to their service. At no point he does, or ever will, have direct access to the AWS console. The tools mentioned are the full-time job of dozens of other engineers (and PMs, EMs and managers). This process takes days to weeks to complete.
Bootstrapped companies generally don't do this btw. This is a symptom of venture backed companies.
> I also never seen an eng org where substantial part of it didn’t do useless projects that never amount to anything
Your description applies to a substantial number of business units in that company. They also had a "research institute" whose best result in the last decade was an inaccurate linear regression (not a euphemism for ML).
Name one of business (tech or non tech) where this is ok/accepted and competitive in capitalism.
How long will we keep making these inflated salaries while being known for being wasteful, globally speaking?
Not that we'd need them as we wouldn't have to write as much HCL.
Would've saved another ~30% for minimal difference in performance.
For me this doesn't look like a sensible move especially since with AWS EKS you have a managed, highly-available, multi-AZ control plane.
Unless their product is pretty static and not seeing much development, they're probably in the negative.
What's capex vs opex now? Thats 150k of depreciable assets, probably ones that will be available for use long after all the current staff depart.
Everyone forgets what WhatsApp did with few engineers and less hardware, there's probably more than enough room for them to grow, and they have space to increase capacity.
The cloud has a place, but candidly so does a Datacenter and ownership.
Now you're maintaining two tiers. That's more work, not less.
I think the real question is what's the cost to buy an equal number of training hours so you can pretend your resources are that competent.
Imagine a military that never fought, never did significant exercises and probably doesn't even have cleaning exercises any more on half its inventory.. That's basically how I view a company that had organic IT growth over a few years and hasn't done a major transition in anyone's recent memory.
The move does cost money, once. Then the savings over years add up to a lot. We made this change more than 10 years ago and it was one of the best decisions we ever made.
Recently I've been moving most projects to Hetzner Cloud, it's a pleasure to work with and pleasantly inexpensive. It's a pity they didn't start it 10 years earlier.
Why would you spread FUD? They have several datacenters in different locations, and even if they were as incompetent as OVH (they are not)[0], the destruction of one datacenter doesn't mean you will lose data stored in the remaining ones.
[0] I bet OVH is also way smarter than they were before the fire.
During that time, one of my servers was hacked once (I was stupid enough to start digging Monero on the same system I had some other services installed) and another time one of my users had a weak password and his account was sending spam. In both cases they notified me and gave me the time to fix the problem. I also appreciate human contact and quick replies.
Also, there are companies that manage spot price allocation for you, so you should essentially always pay spot+small_x% and never actually get terminated.
Let's say you're really pushing the connection and your p95 is 900megabit up. That is $200 at colo vs ~$8200 for amazon.
Wait, is this accurate?
If so I need to sign our company up for a savings plan... now. We use RI's but I thought savings plan only applied to instance cost and not bandwidth (and definitely not S3)
You're not saving anything doing it yourself.
And you've just given yourself the massive inconvenience of running a HA Kubernetes control plane.