Intel Gets Serious About Neuromorphic, Cognitive Computing Future
nextplatform.com
nextplatform.com
Meanwhile, upcoming (non-neuromorphic) AI processors are taking two directions: larger numbers of simplified GPU-type cores (such as NVIDIA Xavier and Intel's Lake Crest/Nervana chips), and FPGAs.
Simplifying cores means lower precision, as fp32 and fp64 are overkill for neural networks and take up lots of silicon. The current NVIDIA Pascal added fp16 and byte operations such as the DP4A convolution instruction[1]. Even smaller precision is practical (down to 1 bit with XNORnet[2], and the DoReFa paper[3] gives an excellent summary of the falloff in accuracy through 32-8-4-2-1 bits for weights, activations, and gradients).
[1] https://devblogs.nvidia.com/parallelforall/mixed-precision-p...
[2] XNORnet, https://arxiv.org/abs/1603.05279
[3] DoReFa, https://arxiv.org/abs/1606.06160
[1] "We propose DoReFa-Net, a method to train convolutional neural networks that have low bitwidth weights and activations using low bitwidth parameter gradients. In particular, during backward pass, parameter gradients are stochastically quantized to low bitwidth numbers before being propagated to convolutional layers" (from the abstract)
Neuromorphic designs like IBM's True North are more hardwired, and that's the limitation towards general purpose use. Yann LeCun's remarks on True North:
LeCun's primary critique is that binary wont work; "to get good results on a task like ImageNet you need about 8 bit of precision on the neuron states". There was no evidence for his claim then, and now this is clearly false [2, 3].
LeCun's post is based more in pride than reason; he spends most of the time talking about NeuFlow which at one point was a competitor to True North for funding. In the end, NeuFlow never became a chip, but True North did.
[0] https://papers.nips.cc/paper/5862-backpropagation-for-energy...
[1] https://arxiv.org/pdf/1603.08270.pdf
1) Nothing on imagenet
2) They already fall to 83% accuracy on CIFAR-10! Imagine how bad imagenet will be! If they string many chips together (exploding their power consumption, since here comes the Von Neumann Bottleneck of data movement), they get a paltry 89%.....
Meanwhile even squeezenet achieves better results.
I dont think you understand their architecture, and neither did LeCun. The Von Neumann Bottleneck is a specific term referring to limited throughput between data in memory and compute in the CPU. TrueNorth is not a Von Neumann architecture and does not have this bottleneck. Memory is located adjacent to compute elements in True North. For comparison, GPU's have very tiny amounts of on chip memory, and have to spend lots of energy copying data back and forth to off chip memory, which is why they are investing heavily in approaches like HBM. FPGAs also dont have as much memory as an ASIC because they need to dedicate space to reprogrammable logic, integrated ARM cores/DSPs, etc.
The chips can be laid out in flexible topologies such as a grid. While it's true that communication between chips is more power intensive than within a chip, this cost is only occurred for the relatively small amount of traffic sent between chips versus computed locally. Hierarchy and small world nature of neural networks can mean that there is more local computation than you would expect naively, and a grid can mean a spike routed from one core to the furthest core would be O(sqrt(N)) instead of O(N).
Pretty sure LeCun and I understand what the Von Neumann Bottleneck is thank you very much.
The thing is though that TrueNorth isn't doing anything special by pouring a ton of memory on die, and even in GPUs, CNN runtimes and energy consumption is dominated by compute.
LeCun never said anything about the Von Neumann Bottleneck. TrueNorth is not a Von Neumann architecture; it does not have a memory bus; it does not have the Von Neumann Bottleneck [0,2,3]. From wikipedia [1]:
"TrueNorth circumvents the von-Neumann-architecture bottlenecks and is very energy-efficient, consuming 70 milliwatts, about 1/10,000th the power density of conventional microprocessors"
If you disagree, please explain how you think the Von Neumann Bottleneck applies here.
With regards to energy consumption keep in mind the smallest GPUs (TX1) are ~10W, typical FGPAs ~1W, versus 70mW for TrueNorth! It's popular to hate on TrueNorth but you could throw 10 of them together and still be fantastically more efficient than anything else today - that's super cool to me! It required lots of special engineering effort to get right, such as building a lot of on chip memory.
On chip memory is one of the most difficult components to get right, minimizing transistors while not breaking physics. It's not as simple as "pouring tons of memory on a die" and requires specialized engineers that hand-layout these components. The event driven asynchronous nature of TrueNorth is fairly unique and undoubtedly added complexity to the memory design.
Do you have any references or evidence for CNN runtimes being mostly dominated by compute? The operations performed in a CNN are more than just convolution; for every input you multiply by a weight, you now have a memory bound problem, which is much more expensive than ALU operations. Don't just take my word for it, listen to Bill Dally (Chief Scientist at NVIDIA, Stanford CS prof, and general computer architecture badass) [4]:
"State-of-the-art deep neural networks (DNNs) have hundreds of millions of connections and are both computationally and memory intensive, making them difficult to deploy on embedded systems with limited hardware resources and power budgets. While custom hardware helps the computation, fetching weights from DRAM is two orders of magnitude more expensive than ALU operations, and dominates the required power."
This is what TrueNorth got right, and made its bet completing design before AlexNet was published. That was a time where Hinton was viewed by the ML community as a heretic talking about RBMs and backprop and hardly anyone believed him. TrueNorth, like NNs at the time, gets some shade by doing things differently that over time we're seeing validated by other researchers and architectures incorporating.
I recommend reading [4] if you haven't already, as it is rich in insights for building efficient NN architectures.
[0] https://en.wikipedia.org/wiki/Von_Neumann_architecture#Von_N...
[1] https://en.wikipedia.org/wiki/TrueNorth
[2] http://ieeexplore.ieee.org/document/7229264/?reload=true&arn...
Anyways, as for showing that convolution runtimes are dominated by compute, not lookup, as much as I tried, I couldn't find the goddamn chart that shows the breakdown, but I did see that chart somewhere and my own experiments show that to be true. It IS true however, that in general memory access is far more expensive than the operations, but deep nets are basically "take the data in and chew on it for a long time". Besides, the way to beat the Von Neumann Bottleneck likely lies in fabrication technology, not design (like HBM2 and TSVs). What makes you think their SRAM cell is custom? It appears to be a standard SRAM cell. That's what I meant by "pouring memory on die".
And the von neumann bottleneck is primarily caused by memory access (aka data movement) being expensive. What happens if you have to move data between multiple truenorth chips?
The term was originally focused on both data and program memory being on the other side of a shared bus from the CPU, which meant you could only do one at a time. If you were trying to figure out your next instruction, you wait. If you then need to data, you wait. It wasn't a problem back with slowly executing EDVAC code [0]. Based on this definition most architectures today do not have the Von Neumann Bottleneck as they are not Von Neumann architectures [1, 2, 3].
A slightly looser definition of the Von Neumann Bottleneck refers to the separation between CPU and memory with a single bus. This likely originated because fully Von Neumann architectures are so rare but that the general problem is similar enough to share the name. GPUs dont have this issue because they employ parallelism thru multiple memory ports talking to off chip RAM. TrueNorth also doesn't have this issue because it has 4096 parallel cores with their own localized memory and no off chip memory. There could certainly be other bottlenecks in the system, even with the memory system, but those wouldn't be the Von Neumann Bottleneck [0].
[0] https://en.wikipedia.org/wiki/Von_Neumann_architecture#Von_N...
[1] https://news.ycombinator.com/item?id=2645652
[2] http://ithare.com/modified-harvard-architecture-clarifying-c...
[3] http://cs.stackexchange.com/questions/24599/which-architectu...
By "pour memory on die" I mean that the memory is on die, clearly there are some special techniques being used to manage that memory, but physically, this is what's saving power.
As you can see, ~10-1000X (the scale is logarithmic) more is spent on compute rather than data movement, and that's with DDR, not even HBM2, let alone on-chip!
These numbers are meaningless. If you want to compare power consumption for different chips, you need to make sure they:
1. Perform the same task: running the same algorithm on the same data
2. Use the same precision (number of bits) in both data storage, and computation.
3. Achieve the same accuracy on the benchmark.
4. Run at the same speed (finish the benchmark at the same time). In other words, look at energy per task, not per time.
If even a single one of these conditions is not met, you're comparing apples to oranges. No valid comparisons have been made so far, that I know of.
p.s. The numbers you provided are off even ignoring my main point: typical power consumption of an FPGA chip is 10-40W, and I don't know where you got 70mW for TrueNorth, and what it represents.
However for inference tasks, low precision GEMM (as you said) goes a long way and better than what you often get. That's why chips like Movidius' Myriad are getting popular and are more similar to DSPs than neuromorphic designs.
I agree that Intel's neuromorphic group doesn't get it, but other groups have taken neuromophic design principles that lead to efficient designs. For example, TrueNorth is very low precision, has great data locality, and though it was designed over 5 years ago can still use modern convolutional networks only imagined afterwards [0]. But its silicon implementation is not very brain like.
The article assumes we know, but I haven't heard of it. And the wikipedia article on "neuromorphic engineering" talks about stuff like analog circuits, copying how neurons work, and memristors, none of which seem that related.
Edit, links: Paper: https://docs.google.com/file/d/0B7QHR9a8j1iiU3RxSHZSNFh2cEdv... Slides: https://docs.google.com/file/d/0B7QHR9a8j1iiSE1ET2ZTb09aNFBP...
[1]: Which are wired something like this: http://neuralnetworksanddeeplearning.com/images/tikz40.png - Notice that dense connections and layered architecture. For all intents and purposes, this what neural nets look like today because of how easily it is to treat a NN with this specific wiring as a chain of tensor computations and thus execute on more conventional hardware.
I believe the idea is that if you simulate too many of them, something useful will happen.
One human brain-equivalent NN on classic architecture costs ~$70M and uses ~100 houses worth of power.
In any case though my guess is that a human neuron accomplishes way more than a hidden unit in a NN so it may be fair to view that as a lower-bound.
Instead, the future is recording and uploading your observations to the cloud, with data scientist and neural net wizards training over this dataset on a cluster with tons of GPUs, and then deploying an optimized model to scrappy low precision inference chips.
This is why FPGA based designs will fail to be stunning. Specialized low precision ASICs more similar to DSPs like Movidius' Myriad (in the Phantom drones and Google's Project Tango devices), Google's TPU, upcoming Qualcomm chips, or Nervana's will become increasingly popular.
[1] Neuromorphic Chips https://www.technologyreview.com/s/526506/neuromorphic-chips...
[2] Numenta papers/videos http://numenta.com/papers-videos-and-more/
It would also be good to see a major chip manufacturer or cloud provider that makes its own chips (Google/IBM) get serious about graph processing chips [3,4] and moving beyond floating point [5].
[3] Novel Graph Processor Architecture https://www.ll.mit.edu/publications/journal/pdf/vol20_no1/20...
[4] Novel Graph Processor Architecture, Prototype System, and Results https://arxiv.org/pdf/1607.06541.pdf
[5] Stanford Seminar: Beyond Floating Point: Next Generation Computer Arithmetic https://www.youtube.com/watch?v=aP0Y1uAA-2Y
Don't mean to minimize the work involved, just trying to decipher the marketing speak.
Perhaps the best success of such analog computation efforts are with Neurogrid (full disclosure I worked with this group): https://web.stanford.edu/group/brainsinsilicon/
One thing to keep in mind is that communication is often still in digital spikes, which some may argue is more neuron like than an analog encoding.
The key to circumventing this is very complicated and our "secret sauce".