Nvidia Pascal GPU Architecture to Provide 10X Speedup for Deep Learning Apps
blogs.nvidia.com
blogs.nvidia.com
Half precision floating point is really something that should have been in hardware a long time ago. There are other applications besides deep learning that could benefit from the dynamic range of floating point, but don't need 32 bits. For example, imaging and audio.
But it still doesn't mean it can't be implemented outside the u.s. or in FPGA's.
Even a simulation of this in software could be really interesting.
It might be not enough if you have a lot of filters and doesn't want to care about their ordering/gains etc
If you have something like 16bit exponent and 48bit mantissa (64-bit) that's absurd
Is this a thing? Something that takes a convolutional neural network and spits out a little app?
For example, offline speech recognition in Android phones uses a CNN that was trained on Google's GPU cluster[1]
[1] I can't find the good reference of this, but Slide 31 of this is outdated, but kind of says it (you aren't training a 2.7M parameter CNN on a phone). http://www.cs.nyu.edu/~eugenew/asr13/lecture_14.pdf
For a simple non-NN analogy, think of spam detection: you want to be able to correct misclassifications, such that it won't make them again. This requires more effort than it sounds: it's not enough to just add the one piece of weighted evidence to the corpus and re-run the algorithm, because that won't necessarily make the filter spit out the new correct answer in the case of the original. Just like there's overfitting, there's also underfitting, and the naive training method results in underfitting.
Thus, what you tend to want is something more like a garbage collector: a process constantly running in a background thread, gradually retraining the system. At any given point, the system will answer questions using an MVCC-like point-in-time view of its beliefs, while those beliefs are getting played with and re-evaluated elsewhere.
Also, every time the NN changes its mind, there will be a subset of non-training samples it has already seen, that it will now classify differently than it originally did when it saw them. Continuously going back and amending its judgements on these is usually helpful.
Where was I? Oh, yes, moving a memory chip a few cm makes the data bus faster? I thought that the speed of light (or electrons in this case) was so quick that a few cm wouldn't make any difference?
2. Light travels ~30 cm per 1 GHz clock cycle. That's approaching the point where we have to worry about it, but...
3. The speed of light constrains latency, not bandwidth. GPUs care much more about bandwidth. The advantage of using shorter wires is it allows increasing bandwidth more easily.
This is because every new bit on the bus needs to charge this parasitic capacitor and this takes time and power.
It's probably a differential drive which makes things more complicated in the details, but I think in general the above is true anyway.
This will make all of the stuff in Caffe (C++?), Torch (C) and Theano (Python) faster by inclusion of cuDNN, a low-level CUDA-optimized library of deep neural network primitives.
https://timdettmers.wordpress.com/2014/08/14/which-gpu-for-d...
tl;dr: GTX Titan X = 0.35 GTX 680 = 0.35 AWS GPU instance (g2.2 and g2.8) = 0.33 GTX 960
GTX Titan X = 0.66 GTX 980 = 0.6 GTX 970 = 0.5 GTX Titan = 0.40 GTX 580
also: I was under the impression single precision was fine for most deep learning applications and double precision doesen't even have good support in most libraries but I guess it depends on the use case.
Rather than a speedup factor of 2-3.