In the real world, for baseline load, the big advantage for many large companies isn't price, but the massive lack of alacrity of many inhouse ops teams. If it takes me 3+ months to provision compute for the simplest, lowest demand services (as is the custom in many large companies full of red tape and arguments about who bears costs), letting teams just spin up anything they want and get billed directly is often a winner, even if it's more expensive. Having entire teams waste months before they can put something in prod is a very different kind of expense in itself.
The cloud native approach would be to modify your app so that it can be scaled up and down so you keep a few machines always running, and scale up and down with your traffic so you only run at capacity when you need it.
Sure, sure, you’re about to say something about discounts? Granted, that’s available, but only for commitments starting at one year or longer!
Okay, fine, I actually agree that there are savings available by reducing head count. The entire network and storage teams can be made redundant, for starters. Even considering that DevOps and cloud infra engineers need to be hired at great expense, this can be a net win…
…but isn’t in my experience. Managers are unwilling or unable to make many people redundant at once, so they stick around and find things to do…
… things like reproducing the mess that kept them employed, but in the cloud.
I’m watching this unfold at about a dozen of my large enterprise customers right now.
Got to get back to work and send the fifty seventh email about spinning up a single VM. Got to run that past every team! It’s no rush, it’s only been about fourteen months now…
Anecdotally, when my previous company was looking at costs, cloud unequivocally came out significantly more expensive, and that wasn’t even a large company (only 2,000 or so employees).
I will grant that we did not have globalization problems to solve (but I’d also wager that lots of businesses prematurely “what if” this scenario anyway).
If you neeed 4 CPUs for your peak load for 4 hours per day, and only 1 of them for the other 20 hours a day, you can save by scaling down to 1 cpu for 85% of the day.
It's also extremely uncommon to have loads that spiky.
And when you do, hybrid is often a solution (use a provider that can provide colo or managed servers for your base load and cloud instances for your peaks, or scale across providers).
Even at that, you said yourself that you can use "cloud" to scale into your spikes.
It takes very unusual load patterns for cloud to win on cost. It does occasionally happen, but far less often than developers tend to think.
There many reasons to choose cloud services, but cost is almost never one of them.
Guy asked what is a cloud workload, I responded. Nitpicking every tiny detail doesn't help.
> There many reasons to choose cloud services, but cost is almost never one of them.
It's cheaper to pay me to manage IAM roles for lambdas and ECS instances for 5% of my time than it is to pay someone full-time to manage some sort of VMware or other system. It's easier and cheaper to find someone with experience with AWS who can provide value to the team and product than it is to find someone who can manage and maintain a cobbled together set of scripts to update apps. There are click and go options for deploying major self hosted services like grafana, k8s with secure details that I can use without spending any time (and time == $$$) learning about the developers preferred deployment scheme.
> It's cheaper to pay me to manage IAM roles for lambdas and ECS instances for 5% of my time than it is to pay someone full-time to manage some sort of VMware or other system.
True, but it's a false equivalence, and one I often see used from people unaware of the ease of contracting this out on a fractional basis.
I used to make a living cleaning up after people who thought cloud was easier, who ended up often spending a fortune untangling years of accumulated cruft that just never happened for my non-cloud customers.
Because you are paying someone else for them.
This is considered rational because those operators are presumably more productive in a pool of people using similar skills to support many customers rather than just one. It is similar to hiring a cleaning service rather than employing individual cleaners in a department of cleaning because cleaning things is not a core competency of business.
It might be less irrational if some amount of compute is part of the core competency of the business. Since "software is eating the world," compute is a core competency of all businesses except for the ones that don't realize it yet.
I've not really seen this work out well. I think it might be true for simple set-ups, letting a tiny developer team also handle infra and support without going nuts doing it, if they set it up that way from the beginning, but more-complex setups always seem to have so damn many sharp edges and moving pieces that support ends up looking similar to what a far more DIY approach (short of building one's own datacenter outright) would, in terms of time lost to it.
... and so does downtime, for that matter.
I'm used to organizations moving out of the cloud when they realize that it's more expensive if you don't have very peaky load demands.
And it's difficult to make that as expensive as a cloud deployment.