(Author here)
I agree with a lot of your points, especially the fact that TCO is practically impossible to calculate, is highly subjective, and should ideally include opportunity costs which, almost by definition vary from person to person.
I've definitely thought through some of your points. First I want to say that if your load is spiky you have no business buying hardware. No matter how you slice it, you're simply not going to be able to get hundreds of GPUs of throughput with an 8 GPU machine so it's not possible to get the same "product". So for those doing short bursts of hyperparameter search, stick with cloud.
But, like you said, for those with high utilization and known utilization patterns, it makes a lot of sense to go on-prem. Let's just say pretty much anybody who is buying a reserved instance from AWS should consider buying hardware instead.
>- Cost associated with finding an admin who understands how this thing works
Pricing based on what co location facilities will provide you without much search and is pretty generous. $10k per year per server is silly high and should cover that search cost.
> - Try before you buy
You can try essentially this exact machine on AWS. That instance and hardware are almost exactly the same. Now of course this doesn't apply to many instance types.
> - Time it takes for the box to be built, shipped and sent to the data center
5 business days is our mean lead time for that unit. Most workstations are 2 days.
> - Time (and cost) it takes to install software, drivers, etc
It comes with that stuff pre-installed as mentioned in the article. See Lambda Stack for driver / framework / CUDA woes: https://lambdalabs.com/lambda-stack-deep-learning-software
> - 3-year? What is the useful life of this? When does it seemingly become obsolete?
I'll admit it's a long time but 3 years is about right for GPUs from my experience. I personally purchased new GPUs for my workstations in 2012 (Fermi), 2015 (Maxwell), 2017 (Pascal).
> - AWS will continually upgrade their hardware and you keep paying the same
Yes, but only if you never commit to a reserved instance, in which case costs are 3x higher. If you buy a reserved instance you don't get upgraded hardware.
> - Spending $90k instead of $184k in year 1 with the option to turn it off if you want (no longer need). This could be very valuable for a startup who wants elastic spending patterns.
Yea, most of these are people who are already using GPUs really often for internal training infra and are finding it expensive.
> - Returns, breakage, warranty in case of a hardware failure
Our hardware comes with a 3 year parts warranty. Of course, this doesn't pay for your time lost but it's pretty rare to see parts fail and the $10k / year more than covers in-colo swap outs.
I agree with all of your points. I don't think that we really disagree with much here. TCO is hard to calculate :).