If TPU teraflops are reported at fp16, then the number would be half.
Titan X offers 12 TFLOPs per GPU https://blogs.nvidia.com/blog/2017/04/06/titan-xp/
If TPU teraflops are reported at fp16, then the number would be half.
Titan X offers 12 TFLOPs per GPU https://blogs.nvidia.com/blog/2017/04/06/titan-xp/
Where are you getting the extra 3 zeros from?
(Also 'up to 180 TFLOPs' is a bit misleading. As I read in some benchmarking posts by Google as well as by nVidida, a TPU is much faster than a GPU for doing inference, but they haven't released any data on training performance, the real bottleneck IMO).
Might also be TPU gen2 given the "64 GB of ultra-high-bandwidth memory" was not something IIRC the recent TPU paper was talking about.
And this is here now, while Volta is paper-launched and will be very scarce until late 2017/2018.
A GPU is also comprised of multiple chips (RAM, etc). I don't think "performance per discrete piece of silicon" is an interesting metric.
The closest comparison in terms of size would be 1 Volta DGX-1 (8x V100s) compared to 2 TPU modules (8x TPU2 chips).
Volta DGX-1: https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Cent...
TPU Module: https://storage.googleapis.com/gweb-uniblog-publish-prod/ima...
And for completeness, this is the size of a single V100: https://cdn.arstechnica.net/wp-content/uploads/2017/05/34446...
You can see that 8x V100s are still more computationally dense than 8x TPU2 chips. Density is a very important factor in datacenter design.
http://hothardware.com/ContentImages/NewsItem/37034/content/...
(That's a P100 server, obviously, but it's the same kind of design. The P100 module looks a lot like the V100 module.)
The point of my previous comment is that it doesn't make sense to compare an entire TPU module (containing 4x TPUs) to a V100 board.
The closest comparison to a TPU module (with 4x TPUs) would be a dgx-1 board, which contains the nvlink bus that you mention, but also contains 8x V100 boards, hence why in my previous post I said you should compare the compute performance of a Volta DGX-1 (8x V100s) to 2 TPU modules (8x TPUs).
At the end of the day, it is simple, in a given area of space, you can get more compute performance from provisioning that area with V100s (in the form of using dgx-1s) versus provisioning it with TPUs (in the form of using TPU modules).
That's true if you're putting the chips in your own data center, because density affects TCO.
But assuming that Google does not sell TPU hardware, this isn't an important metric to anyone other than Google. The question that is important is, what does the shape of the curve plotting "dollars spent" versus "time spent waiting for my model to train" look like?
The shape of that curve is affected by TCO, and TCO is certainly affected by density. But there's a lot more to it than that.
But not to a single client.
Similarly, Google is not offering a 1000 TPU cluster to each client.
Your metric is strange to say the least.
The previous TPU was aimed at inference. The Cloud TPUs do training. But you're right - I haven't seen any publications about training performance yet. But they'll be available in cloud (currently in alpha), and once they go GA, I'd expect to see a raft of benchmarks by third parties. I can't wait. :)
V100 is 120 teraflops of tensor ops per chip: https://arstechnica.com/gadgets/2017/05/nvidia-tesla-v100-gp...
The closest comparison in terms of size would be 1 Volta DGX-1 (8x V100s) compared to 2 TPU modules (8x TPU2 chips).
Volta DGX-1: https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Cent...
TPU Module: https://storage.googleapis.com/gweb-uniblog-publish-prod/ima...
And for completeness, this is the size of a single V100: https://cdn.arstechnica.net/wp-content/uploads/2017/05/34446...
You can see that 8x V100s are still more computationally dense than 8x TPU2 chips. Density is a very important factor in datacenter design.