I feel like thats becoming tensorflows 'native platform'...
I feel like thats becoming tensorflows 'native platform'...
TPUv2 were benchmarked against NVIDIa's K80 at 30 times the performance and had a peak of 92 TOPS [2] while the 2080 TI is at 440 TOPS [3] [1] https://www.anandtech.com/show/11749/hot-chips-google-tpu-pe... [2] https://www.tomshardware.com/news/tpu-v2-google-machine-lear... [3] https://www.microway.com/knowledge-center-articles/compariso...
nVidia is destroying the TPUs right now and Google is desperate to keep their perception in the public eye as the king of AI (which tbf, they probably are, compute capabilities aside)
In this case the data was from the time Google introduced the TPU internally, when the K80 was very much up-to-date. It also makes sense because the K80 was the only GPU offered in GCP.
Also, there's no need for assumptions when you know what's going on in the design team.
(disclaimer: while I'm part-time at Google, this is my personal impression, not an official statement, etc., etc.)
(420 TFlop/s, 128GB HBM)
The price/performance ratio of rented TPUv2 or V100 can't match the price/performance ratio of owning the system if you are doing lots of learning/inference.
If the model fits inside 2080 Ti and the work is not tightly time restricted, 2080 Ti (the whole $2.5k system) should be more economic choice after six months or less (full utilization 24/7).
DAWNBench does benchmarks.
"At the time DAWNBench contest closed on April 2018, the lowest training cost by non-TPU processors was $72.40 (for training ResNet-50 at 93% accuracy with ImageNet using spot instance). With Cloud TPU v2 pre-emptible pricing, you can finish the same training at $12.87. It's less than 1/5th of non-TPU cost. "
Here is a link to DAWNBench.