Emergent Chip Vastly Accelerates Deep Neural Networks
nextplatform.com
nextplatform.com
The compressed network achieves a decent speedup and energy saving on current hardware too (desktop/mobile) without significant loss of accuracy
The chip is able to fit much more into SRAM because the chip uses network compression, pruning, quantization, Huffman encoding, etc.
The term static differentiates SRAM from DRAM (dynamic random-access memory) which must be periodically refreshed. SRAM is faster and more expensive than DRAM; it is typically used for CPU cache while DRAM is used for a computer's main memory.
WP
This may lead to high quality speech recognition or image recognition right on the device, without requiring a beast of a GPU.
It's particularly bad if you want to run them on a small, battery powered device. Far more so if you only have one image to process, here they're seeing single image speedups of over 10x compared to GPUs (but slower than batched processing) and energy efficiencies about 1000+ times better than mobile GPUs (and far more compared to the beasts in your desktop).
This is absolutely incorrect.
A mobile device can execute a pretrained model fine. See, for example [1][2][3]. Google's new TensorFlow NN system is explicitly designed to be able to run on mobile devices and comes with a pre-trained image classification NN that workd fine on mobile devices.
This doesn't mean that energy saving is unimportant on a mobile device of course. But there are very widely deployed production systems (eg all of Android) that use them now with no special GPU acceleration.
[1] http://googleresearch.blogspot.com.au/2015/07/how-google-tra...
[2] http://googleresearch.blogspot.com.au/2015/09/google-voice-s...
[3] http://static.googleusercontent.com/media/research.google.co...
https://github.com/tensorflow/tensorflow/tree/master/tensorf...
Of course it can. The problem is with the size of the network you might want to run.
From your first link, which is about single character level image processing
> We needed to develop a very small neural net, and put severe limits on how much we tried to teach it—in essence, put an upper bound on the density of information it handles.
And from the voice training paper:
> While our server-based model has 50M parameters (k = 4, nh = 2560, ni = 26 and no = 7969), to reduce the memory and computation requirement for the embedded model, we experimented with a variety of sizes and chose k = 6, nh = 512, ni = 16 and no = 2000, or 2.7M parameters
AlexNet is, what, 60M+? VGG is pretty big too.
I guess I wasn't too clear. Yes, there are good results from nets that we can run on mobile devices in realtime. We do, however, want to run significantly larger nets.
I don't know how many parameters Inception v3 has, but I know Google considers it more efficient than VGG ("Although our network is 42 layers deep, our computation cost is only about 2.5 higher than that of GoogLeNet and it is still much more efficient than VGGNet".)
Yes, being able to run big networks is great. But ultimately it's what you do with it, and a 3% error rate on ImageNet is a pretty compelling argument that size isn't the only factor.
[1] https://github.com/tensorflow/tensorflow/tree/master/tensorf...
You can do limited tasks but not the kind of things we can do on desktop
So, you might still send the request to the network to continue training the model, but by the time you do, your answer has already computed on the local machine for local consumption.
Combining these two techniques would be really cool and if the bitwise network can work with larger, more complex networks like VGG would be a massive game-changer, allowing these nets to fit on almost any device.
Is that no longer the case? ARM architectures seem to be beating x86 for low power, mobile devices. GPUs are being used for many easily parallelizable workloads.
Is an overall slow down in Moore's law making chips designed for specific tasks (like Deep Neural Nets), attractive again?
Underneath, Intel processors translate x86 to simpler RISC like microcode. They could make a more efficient chip without this translation component, and probably holds them back a bit in low power stuff.
Now that process shrinking has become extremely difficult, everybody else is starting to catch up.
What is power?