Cloud TPU Pods Break AI Training Records
cloud.google.com
cloud.google.com
Seriously.
Transformer: 1024 TPUs are twice as fast as 480 GPUs.
Resnet50: 1536 GPUs are about as fast as 1024 TPUs.
SSD: 1024 TPUs are twice as fast as 240 GPUs.
Great.
I know. Within MLperf, to reach consensus, it was agreed that “dollars are great, but also game-able via price cuts, lets normalize to a reference system”. It’s not the largest box though, it’s one that everyone was able to provide, so that it lets you compare through one level of indirection.
I don’t agree with it, because it makes it so hard to compare to funky hardware (“Should I do inference on this chip from < startup >? I dunno, what’s it going to cost?”). Having to keep going back to some reference box is undesirable, but there’s nothing stopping someone from making a spreadsheet equivalent that translates these to dollars on the cloud providers (harder for on-prem submissions which reopens the “rent vs buy” debate).
I'm objecting to the fact that your marketing team (at least, I hope it was your marketing team) chose to represent 5 different GPU systems as the same color and 3 different TPU systems as another color in one bar chart.
There's a book with a chapter or two on how to reasonably compare the performance of computer systems. Given that the author (Patterson) works on TPUs, I would expect better.
Our impression is that top-line performance comparisons independent of system size (like the comparison in the blog post) and performance-per-dollar comparisons are the easiest to understand, but we're certainly open to other ideas.
Then do that, just don't do it all in one chart with the systems only specified in the fine print.
Cloud TPUs are designed to maximize performance-per-dollar, so you are right that pure performance comparisons at maximum scale don't tell the whole story.
The most straightforward performance-per-dollar comparisons would be among several different hardware configurations across the major public clouds. However, we haven't yet seen any other public cloud MLPerf submissions at scales comparable to Cloud TPU Pods, so there isn't currently a strong baseline available for comparison. It's also not clear whether public cloud networking will ultimately be able to match the performance of the network hardware that was used to produce the largest-scale on-premise MLPerf submissions.
The best apples-to-apples performance-per-dollar comparison we have publicly available was published last fall, and it compared the performance and cost of using various Cloud TPU v2 Pod slice sizes with the performance and cost of using various numbers of V100 GPUs attached to a single GCP host:
https://cloud.google.com/blog/products/ai-machine-learning/n...
We went to great lengths to ensure that we trained exactly the same version of ResNet-50 to the same accuracy in the same way across all hardware configurations. The methodology predated MLPerf and is documented in full here:
https://github.com/tensorflow/tpu/blob/master/benchmarks/Res...
If you were going to do a similar performance-per-dollar comparison today, the simplest approach might be to try to get the code from NVIDIA's MLPerf 0.6 submissions running at scale on one or more major public clouds using the fastest-available networking technology that each cloud provides:
https://github.com/mlperf/training_results_v0.6/tree/master/...
It would be very interesting to see how distributed training performance using large-scale GPU clusters in public clouds compares with the published on-premise MLPerf performance numbers using exactly the same MLPerf code and methodology. With these measurements in hand, it would then be straightforward to make performance-per-dollar comparisons with Cloud TPU v3 Pod slices of various sizes.
You are talking to your customers who is paying or considering paying, or in search of products.
Take the feedback, if it can be done, and it's beneficial, do it and report so.
Or stop explaining... That's simply not professional for a cloud provider...
Performance per watt doesn't care where the systems are (on-prem vs cloud) and should track perf/$ (unless perf/$ is mostly determined by subsidies).
For instance, the DGX-2 has an MSRP of $399,000 and consumes 10 kW of power [1]. The average commercial electricity cost across the US is about $0.11/kWh [2], so a DGX-2 running at full tilt costs $1.10 an hour in electricity. Thus, a DGX-2 running at 100% utilization for 3 years costs $9,636 in electricity, which is ~2.5% of the cost of the box itself.
Of course, you probably could get DGX-2s for a lot cheaper if you are buying 100 of them, but the acquisition costs are still going to be significant vis-a-vis power costs.
[1]: https://www.anandtech.com/show/12587/nvidias-dgx2-sixteen-v1...
[1] https://dawn.cs.stanford.edu/benchmark/index.html#imagenet-t...
Unfortunately, their cost numbers aren't actual cost, as they don't allow using AWS spot instance pricing (which is what fast.ai does).
Said another way, spot pricing results aren’t as reproducible in the research sense. It’s also a similar constant factor between the main providers. We felt that once you knew relative price/perf, you can choose to do whatever economic analysis you’d prefer (e.g., maybe if you don’t have any datacenter space yourself, there is no price you’d pay for hardware on-premises, or maybe you are willing to use spot or preemptible).
I'd recommend doing a performance-per-dollar comparison before drawing this conclusion.
It’s hard to compare perf per dollar. Your electricity costs and GPU costs maybe vastly different from mine.
Now the real problem is, can you actually get a TPU pod in practice?
48 to 1 hour would have allowed us to run a few experiment per work day. That would be huge.
If you told me that the next generation of TPUs could run the current 48 hour GPU training set in 1 hour, that would be something.
But 2x faster won’t make a difference. Generally, startups for a new chip need to demonstrate 100x improvement over the general purpose approach to get funding. So this seems like another Google vanity project.
If you really wanted to get from 48 hours to multiple experiments per work day, you could (probably) do that right now by scaling horizontally. However, the cost becomes a major issue when you do that on GPUs. This is where the TPU shines since you can scale out without cost crippling you.
I work on training on GPUs and the TPU is definitely an incredibly useful piece of hardware, not a vanity project.
But, TPUs are not standard, can’t be used for any other usecase. So those savings might not actually be realizable when everything is taken into account.
TPUs do have the memory capacity advantage though (over 2080Ti).
The DGX-2h is a beast! Don’t parse this as “huh, TPU Pods are about the same as just a few V100s”. The data sheet [1] is probably the easiest to follow, but their writeup is more informative [2].
These are souped up V100s, with awesome networking, which is pretty similar in style to a TPU Pod. So I’d say that they’re both purpose built systems for distributed ML training. The name for the NVIDIA system is even “DGX SuperPOD” :).
[1] https://www.nvidia.com/content/dam/en-zz/es_em/Solutions/Dat...
[2] https://devblogs.nvidia.com/dgx-superpod-world-record-superc...
As mentioned in other comments, I'd recommend doing a performance-per-dollar comparison in addition to looking at this pure performance comparison at maximum scale.
> 1024 TPUs are twice as fast as 480 GPUs
Might it be because there are twice as many?
We submitted multiple results using various Cloud TPU v3 Pod slice sizes to show the current achievable Transformer training efficiency at several scales:
https://mlperf.org/training-results-0-6
We're actively improving the whole TPU software stack, so training efficiency is likely to continue to increase over time.
My startup is trying to develop wafer scale integration where you collapse 2 racks of the network, metal boxes, power and cooling into a 300mm wafer immersed a 100 mm x 400mm box with a few fibers and three power cables coming out. That can save around 900% of the capital cost and orders of magnitude of power (especially if you put the box in a building where the waste heat is not wasted but used to heat water for showering and space heating).
As it is now, these datacenter customers and hyperscalers don't seem to care about the enormous waste and cost of paying for inefficient hardware. Considering the enormous cost and carbon emmission savings a wafer scale integration would bring (and the competitive advantage), you would be suprised how hard it is to get funding from them to develop it.
I suspect is also the main reason they don't care to publish normalized benchmarks for $/performance/joule, as it would demonstrate how wasteful it all is.
That's probably because there has been no evidence that wafer scale integration can actually work.