The "NVIDIA Tesla Family Specification Comparison" table indicates 112 TFLOPS "Tensor Performance
(Deep Learning)" for Tesla V100 (PCIe).
Is that double precision?
The Nvidia 1080 Ti has a double precision performance of 332 GFLOPS [1]. If the above number is for double precision computing, the Tesla V100 (PCIe) would be about 337 times as fast (!!)
Does anyone have more insight into these numbers?
[1] https://en.wikipedia.org/wiki/List_of_Nvidia_graphics_proces...
UPDATE : I should have read the article more carefully, it seems to be a mix of FP16 (half precision) and FP32 (single precision). That would likely mean a factor ~10 in computation performance (specifically for deep learning)