Also that is a 8xA100 system as others have noted, but it is the 40GB one which can be found on eBay for as low as $3k if you go with the SXM4 one (although the price of supporting components may vary) or $5k for the PCI-e version.
You can build a lot of stuff on top of these two.
It's absolutely worth the money when you look at the whole picture. Also lambda labs never has availability. I actually can schedule a distributed cluster on AWS.
That highly depends on many things. If you run a business with a relatively steady load that doesn't need to scale quickly multiple times per day, AWS is definitely not for you. Take Let's Encrypt[1] as an example. Just because cloud is the hype doesn't mean it's always worth it.
Edit: Or a personal experience: I had a customer that insisted on building their website on AWS. They weren't expecting high traffic loads and didn't need high availability, so I suggested to just use a VPS for $50 a month. They wanted to go the AWS route. Now their website is super scalable with all the cool buzzwords and it costs them $400 a month to run. Great! And in addition, the whole setup is way more complex to maintain since it's built on AWS instead of just a simple website with a database and some cache.
S3 and EC2 are priced very competitively.
Your pricing only matters if you have availability.
Spending the capex/opex to run a cluster of compute isn't easy or cheap. It isn't just the cost of the GPU, but the cost of everything else around it that isn't just monetary.
Make adoption easy, give a free base tier but charge more could be a very effective model to get start ups stuck on you. It even probably makes adoption by small teams in big companies possible that can then grow ...
Answer these questions, and the equation shifts a bunch!
A full rack with 16 amps usable power and some bandwidth is $400/month in Kansas City, MO. That is enough to power 5x A100s 24x7, so 10k plus $80 per month each, amortized, of course many more A100s would drop the price.
Once installed in the rack ($250 1 time cost) you shouldn't need to touch it. So 10k plus $1250 per A100, per year including power. You can put 2 or 3 A100s per cheapo Celeron based CPU with motherboards.
Of course if doing very bursty work then it may well make sense to rent...
You also left out the data tech costs- probably at least $50K/individual-year in KC (although I guess I'd just work for free ribs).
If you're putting A100s into celeron motherboards... I don't know what to say. You're not saving money by putting a ferrari engine in a prius.
You hire people who are employed by the DC, by the hour for DC work. How much work is there to do once it is screwed into the rack?
The problem though is that getting 2-3MW of power in the US is increasingly difficult and you're going to pay a lot more for it since the cheap stuff is already taken.
Even more distressing is that if you're going to build new data center space, you can't get the rest of the stuff in the supply chain... backup gennies, transformers, cooling towers, etc...
To train a top model you need hundreds of them in a very advanced datacenter.
You can't just plug gpus into standard systems and train, everything is custom.
The technical talent required for these systems is rare to say the least. The technical talent to make a model is also rare.
I trained a few foundation models with images, and I would NEVER buy any of them. These guys are on a wildly different scale than basically everyone.