With physical hardware, the sysadmin will spend a lot of his time fighting the hardware supplier and the colo to get hardware in place and running, then reinvent the tooling. Any trip to the facility is half a day lost.
With physical hardware, the sysadmin will spend a lot of his time fighting the hardware supplier and the colo to get hardware in place and running, then reinvent the tooling. Any trip to the facility is half a day lost.
Have you never lost half a day of getting work done because Amazon's EC2 API was failing to launch new instances? I have! [1]
To simply hand wave away hosting your own hardware as infeasible and more costly (especially when it's been proven to be more costly) is unproductive at best.
[1] https://hn.algolia.com/?query=AWS%20outage&sort=byPopularity... (HN search of "AWS Outage")
Self hosting is doable up to 2 or 3 servers with a very few services. Beyond that it's really suffering from the lack of tooling, the lack of isolation and the lead time of weeks to get anything more in place.
DELL wants to re verify by phone 3 times before they bill or ship anything. Don't ask me why, I don't know, I keep begging them to stop doing that.
I've had the API down about 4 times in a year, always for specific types of instance, if I remember the count well. Never more than a few hours. It's nothing in comparison to the weeks it takes to ship and setup physical hardware.
There are plenty of companies with their own datacenters of hundreds or even thousands of machines. Yes, these companies usually have deep pockets and their own dedicated staff to take care of these servers. This is how it was done before AWS existed, and is likely to continue long after AWS ceases to exist.
You need to get a Dell sales rep you deal with everytime you order. You'll get better pricing too.
About the most complicated tooling we had was Anaconda.
I think I lost three hard drives in that time and all three were hotswapped with zero downtime under warranty. No other failures. And that's not an outlier - I've been doing this 25+ years and the hardware failure rate has been pretty consistently low.
I couldn't imagine having to manage 300 physical servers in a team of two. The hardware, OS and networking alone are more than a full time job with 24/7 oncall expectations. In fact that's why I don't do this anymore.
Pretty sure my current company had more failed hard drives this month than you all these years. I remember a shipment of servers once with more dead motherboards than that.
The point is that managing on-prem equipment is the easiest its ever been. When I started the ratio of admin to box was 1:50, now it's easily 1:300.
The key phrase in your post is 'doing this 25+ years'.
What you're providing is the /confidence/ that you can do it. Most outfits nowadays probably have no one who has that.
At the ultra-small scale (1 person company), it's arguably cheaper to run from a single server you bought.
Mid-sized companies (~2-1000 people) probably benefit the most from cloud computing because they need fewer dedicated salaried employees to manage them.
Beyond ~1000 employees however, you're going to need that expertise regardless of whether you use cloud or on-prem. And the markup of those thousands of cloud machines will start to become significant. You'll need to do a cost/benefit analysis and I can see it going in either direction.
Rather, the cost savings come from the salaries of you and the "one other guy" you mentioned. If cloud hardware can bring your staff of 2 down to 1, then the cost savings are huge. More-so if they can bring it down to 0 and have devs manage the hosted infrastructure. Each of your salaries is likely near or higher than the cost of running all those machines.
Developers and Administrators have different priorities and I can count on one hand the number of true "DevOps" people I know - the rest are either decent admins and poor devs or decent devs but poor admins. The idea is nice but the reality isn't quite what its made out to be.
Physical instances have too much over provisioning and zero flexibility. You can easily be paying for an order of magnitude above the actually used capacity. (This can be helped with VmWare virtualization instead of bare metal but the amount of new companies using VmWare is shrinking and licenses are expensive.)
Whereas with the cloud, you can half any instance CPU/memory to save half the money. There is even build in monitoring to analyze usage. There is little waste.
The cost of the hardware is not the issue.
Paying for 600 machines when you actually need 300 probably "wastes" $50-80k/year in over-provisioned hardware. Hiring a skilled admin to figure out what's going on, optimize the network, and reduce those 600 machines down to 300 probably costs at least $150k/yr in their salary+benefits. The cloud does not decrease the cost of the hardware (it's actually more expensive), but it allows you to reduce the number of admins managing it. That's the cost savings.
Given the above example, what would any reasonable general manager do: hire someone to make things more efficient, or just pay for the over-provisioning?
600 servers at $10k each, that's $6M upfront. Then another $6M within 3-5 years to renew them.
Of course in that example you'd pay someone, even multiple people in fact, to make things more efficient.