Much of what allowed us to implement these savings quickly with a small team was the flexibility afforded by cloud infrastructure. Poor decisions are easy to reverse, but in a bare metal world you better be damn sure what you're doing, which slows down the decision-making process and seriously complicates experimentation. The number of people who know how to build out datacenters at the scale of thousands of machines is vanishingly small.
We'd also need to replace IaaS services like ECS, ELB, ElastiCache, RDS, and DynamoDB. There are certainly off-the-shelf replacements, but we'd need to build-out the expertise within our teams to operate these systems. We're talking roughly a dozen or so engineers working full time for many, many months to get these systems in place from scratch, on top of the even larger effort to design and build out datacenters. I'd much rather plow those cycles into efforts like expanding to multiple regions and improving reliability of internal services. That's a much better return on investment for our customers.
Right now we're in the sweet spot for the cloud. We're way too big to run on trivial amounts of hardware, growing at a rate that makes it difficult to stay ahead of demand in a datacenter-centric world, and too small to justify investing in a scalable hardware build-out.
+1. If your entire environment is steady-state, it may be worth considering a migration to bare metal.
That said, you can expect:
* Worse uptime
* A fair bit of retraining / hiring different skillsets
* A loss of engineering focus as they work on the migration instead of feature expansion
It can make sense economically, but that discussion goes well past "the monthly AWS bill."
Disclaimer: I am NOT affiliated to Google in any way.
You might have to stub out some things using e.g. Cassandra or another open source alternative; just use a Docker or VM image without the production redundancy.
You did admit, that you have no idea how much performance is being dropped on the floor due to not using dedicated hardware. So why not test it?
My belief: you would find you are paying AWS far more than they are actually worth.
Your post seemed to indicate that you hadn't done any calculations for the bare metal server with your own vms scenario. But clearly you have, if only mentally.
http://firstround.com/review/the-three-infrastructure-mistak...
Keep a look on your bill because you don't want to run anything. But take advantage of having no hassle for your first years.
My gut feeling is that you have to get to an extraordinary size to realize any meaningful savings, but that's primarily based on Dropbox's migration off AWS (https://www.wired.com/2016/03/epic-story-dropboxs-exodus-ama...).
The entry cost is £500 per day or $1000 for an engineer.
5 hours on the phone to find the hardware and agree on the order with DELL + an afternoon with customer support because they shipped the servers without hard drives + your project is delayed by an entire week because you don't have the resources => 1 day + 1 day + half a week.
These things would have been 5 minutes on a cloud.
As a for-instance, I was asked this exact question a year or so ago about an on-prem object store instead of S3.
The break-even price to go multi-region was ~15PB on paper; that included DC space, hardware, software (build vs buy was another factored discussion), and the staff to run it.
That assessment was delivered with the caveat that their uptime was not going to approach S3's by any stretch of the imagination-- and their infrastructure outages weren't going to sync with "most of the internet's."
It's a complex topic, and there are many hidden costs...
It depends on what your service looks like: CPU intensive, Memory Intensive, Storage intensive? (In reality some unique mix).
You probably won't see a huge savings year one, as you'll be spinning up a lot of new things and have a fairly large CapEx expenditure. Now if your growth pattern is steady/predictable then you should be able to plan out your hardware buys or do a hybrid solution to handle traffic bursts.
One of the nice things about running your own hardware is that there are some costs that are easier to control. Don't need new hardware? Don't need to spend on new hardware for example.
You also have much more control over your environment so you are able to really optimize your code, and infrastructure so that you don't need to scale as large system wise.
But, back to the question on how to model it? You just gotta dig in, and make some educated guesses about performance,test and repeat.