Nvidia Announces Tesla P40 and P4 - Neural Network Inference, Big and Small
anandtech.com
anandtech.com
In case it's not clear from the title (which used to read "...47 INT TOPS"), that's 47 [8-bit integer] tera-operations-per-second. Anandtech says it will "... offer a major boost in inferencing performance, the kind of performance boost in a single generation that we rarely see in the first place, and likely won’t see again." No kidding!
Which isn't to say that the Deep Learning stacks won't get there eventually, but at the moment it's not as easy as flipping a switch.
INT8 is a 4x for inference, but most people aren't using GPUs for inference atm.
https://blogs.aws.amazon.com/bigdata/post/TxGEL8IJ0CAXTK/Gen...
Unrelated question though: any chance you will do blog post/paper about how DSSTNE does automatic model parallelism and gets good sparse performance compared to cuSparse/etc?
And given Frank Seide et al. demonstrated 1-bit SGD in 2014 (https://www.microsoft.com/en-us/research/publication/1-bit-s...) the race to the bottom is just beginning...
Nvidia needs a "Ludicrous Mode" overclock setting for these cards. Push the button for 2.5 seconds of super-high frame rate! With some cool down time required.
This actually is a thing with GPU Boost 3.0. The card boosts up automatically if the temperatures are low, though lately there were issues with the GPU clock frequency oscillating because there wasn't enough hysteresis.
(This, by the way, means that GPUs are thermally constrained rather than timing-constrained.)
Then I saw that it was Nvidia. :)
EDIT... apparently when I last googled years ago this was not the case. I wish I could delete this comment :(
As cool as these cards are, I really hope they become available on AWS soon. The current AWS GPU instances are so weak we're contemplating buying a physical desktop setup.
Currently, P100 wastes about half its area on dedicated FP64 cores, which no one needs for deep learning.