Granted it's probably not at this scale, but it gives you access to a ton of resources.
This makes a huge difference when using more than a single machine. I've done the math and purchased the machines at a previous company- assuming you aren't leaving the machines idle most of the time you save a considerable amount of money and get a lot better performance when building your own cluster.
I mentioned in another comment that the GPU generation is roughly three years between major architecture upgrades- this has held true for a bit now, and that time may even stretch out a little. When the average company builds one of these clusters it's safe to assume they'll either run it for three years or sell it back for some return.
Going with the cloud and assuming you don't commit to several years (losing that adaptability) the yearly cost of a p4 is $287,087. Over three years that's $861,261 to run a single machine. For about $450k you can build out a solid two machine (16gpu) cluster (including infiniband networking gear and a solid NAS) that will easily last three years. There are datacenters which specialize in this and companies that can manage these machines. If you don't have the cash up front you can lease them on good terms and your yearly bill will still be much lower than AWS.
Model training is basically the one use case where I'm really willing to purchase equipment instead of using the cloud. The money it saves is enough to hire one or two more staff members, and the maintenance is shockingly low if you get it setup right to start.
The CPU/GPU part of training is only part of the much larger ecosystem, which is what you're paying to be in when you use AWS.
> you save a considerable amount of money and get a lot better performance when building your own cluster.
This heavily depends on how much benefit you get from improving GPU performance from each generation. A lot of people assume 3/4-yr TCO. If you instead "rent" for 1 year at a time, you've been getting >2x benefits per generation lately.
Most folks also measure "occupancy" for clusters like this rather than "utilization". That is, if a job is using 128 "GPUs" that counts as 128 in use. But that ignores that many jobs might have been just fine with T4s (which are a lot cheaper) versus A100s. (Depends a lot on the model, the I/O, etc.) Once you've bought a physical cluster, you're kind of stuck with that configuration (for better or worse).
tl;dr: It's not just about "idle".
At the same time it's hard to understate how different the prices are here. Our break even point for using on prem compared to AWS was about nine months. After that we saved money for the rest of that hardwares lifetime.
I definitely agree that people shouldn't just rush out and buy these without benchmarking and examining their usecase. The cloud is really good for this! At the same time though I have yet to see any cloud provider with anything even approaching the interlink I can get on prem and that means, so it's basically impossible to get the same performance out of the cloud as it is on prem right now.
At list price or a moderate discount. The folks at this scale aren't paying that :).
I always wonder about performance on these clusters. Back in MY day, I'd wait a week or more for results from my jobs, and immediately resubmit to wait two weeks in a queue for another week of runtime and do lots of data processing in the downtime. Then, I moved to cloud and decided on "what can I afford to do overnight" (IE, I set my time to result to be about 12 hours). I have a hard time justifying additional hardware to get results in 10 minutes versus a day, it seems like at that point you're just using it to get fast cycle times on new ideas, but who has new ideas every 10 minutes?
This is the beauty of the new nvswitch chips and the infiniband networks instead of ethernet. Anyone who is doing this is setting up a fully switched high bandwidth infiniband network with ridiculous traffic between them. Nvidia purchased Mellanox a year or two ago- combine that with the ridiculously awesome nvswitch in the A100 dgx machines and there's a huge jump in cross chip traffic ability. At the same time though a decent mellanox router is probably going to set you back $30k.
Better to find algorithms that need less communication, than to make faster computers that allow you to write algorithms that needs lots of communication. Otherwise you'll always pay $$$ to reach peak GPU performance.
It is really peanuts for US government or EU.