Yann LeCun on the IBM neural net chip
facebook.com
facebook.com
It is not clear how LeCun would propose to send information in his convolutional neural networks in specialized hardware (from what I can tell, he uses FPGAs which are very high power in comparison). In the worst case, that approach sends data between neurons on each time step, which is very inefficient in power. If they do something more clever, I would bet it would start to like sending spikes.
It also seems very short sighted to say that if the hardware is not specifically designed for convolutional neural networks, then it is not the right architecture. It seems like the IBM chip does support convolutional networks, but it might require a few extra time steps to average.
Biology has evolved to use spikes (across many different species). Perhaps evolution didn't get the memo that non-spiking convolutional neural networks is the only architecture worth building. Maybe it will take a few more thousand generations before evolution catches up, but until then spiking neuron architectures seem like a decent gambit ...
Those different species didn't independently arrive at their own neuronal implementation, we're all using the same basic pattern (with heavy modifications).
More importantly though, in biology, most cells use temporal encoding - one spike isn't meant to be a binary event. Instead, the frequency of pulses over time is used to encode intensity. The receiving end can then continuously integrate the signal over time. I'm not sure the IBM chip uses spikes in the same way.
> Perhaps evolution didn't get the memo that non-spiking convolutional neural networks is the only architecture worth building.
Temporal encoding in nature is used for the obvious reason of hardware efficiency. If every axon had 8+ binary signal lines it would become unreasonably complex and large, and prone to failure. This somewhat mirrors the design decision in the TrueNorth chip, again for obvious reasons.
However, the two diverge algorithmically. In nature, signal processing is used to transport relative values over the network. The chip on the other hand really seems to use boolean encoding.
You're right about one point though: it's fairly certain that a large number of configurations and algorithms do produce "working" neural networks. Nobody said that imitating our own implementation details will lead to better results than anything else we might want to try.
And in fact, that's the reason why so many computational implementations of neuronal networks do not use temporal encoding, because encoding values in floating points or even integers comes more naturally to computers than it does to wetware. But it is an abstraction that takes up comparatively many resources.
In this light, and although CNNs aren't the only architecture, his criticisms may be a little more reasonable.
That is great that people want general purpose machine learning chips, the question is how to do it in low power. My guess is that the right architecture will be a mix of ML primitives mixed in with things like spikes (and perhaps other primitives seen found in biology).
Current computers are limited by the amount of power required to process large volumes of data. In contrast, biological neural systems, such as the brain, process large volumes of information in complex ways while consuming very little power. Power savings are achieved in neural systems by the sparse utilizations of hardware resources in time and space. Since many real-world problems are power limited and must process large volumes of data, neuromorphic computers have significant promise.
That may or may not be a good hypothesis, but it seems interesting to investigate. In any case, LeCun's real beef is with the DARPA program managers: he thinks a different area of ANN research would've been a better allocation of funds, because in his view this is not among the most promising lines of research. Not an uncommon reaction to DARPA choices, and not always wrong either, but DARPA's got the money.
[1] http://www.darpa.mil/Our_Work/DSO/Programs/Systems_of_Neurom...
I think this is representative of a lot of AI now. This chip doesn't obviously improve the state of the art on an arbitrary (but standard) benchmark, so LeCun dismisses it. That type of attitude strikes me as over fitting for an arbitrary benchmark (a local optimum) and missing the bigger picture. This chip (and line of research) could help identify how the brain works (which may very well also unlock strong AI).
It would make even less sense for them to have just said, 'oh no, we're busy focusing on other things at the moment' when DARPA tried to start throwing money at them.
Surely you can do this in the spatial domain instead of time? That is, each neuron is one bit, so you have 8 each computing one of the 8 bits he says are necessary? Perhaps the problem is using that value afterwards, I guess.
E.g. you have a 4x4 pixel image you want to feed in to the net and recognize. You subdivide it into 2x2 images (of which you'll have 9), and you send each of those 2x2 images to each neuron. Then the output of the neurons is one if they see waldo, zero if they don't. Why not send everything to every neuron ? That won't work, and they don't have enough inputs anyway in real image sizes.
Actually you'd send more than 9. You'd also include a 2x2 that only includes the corner pixels, to achieve scale independence (recognizing a car whether it's 5x5 pixels or 50x50). Then you'd send that 36 times, but each time rotated by, say 10 degrees. That's how human vision works ("how God does it") and, well, that's how AI is trying to do it. In humans it's not a full haar cascade : we can only see small features using the centre of the retina, and only large features using the rest, and we only allow for limited rotation (meaning humans brains rotate the source image a limited number of times along an exponential curve that only goes up to ~40 degrees rotation).
(This is comparable to a haar cascade)
Now spiking neural networks have suboptimal performance, true, but they have a major advantage : they do unsupervised learning only. You show them a world, and they will build their own model of the world (which isn't as good as our state-of-the-art models for known "worlds" like ImageNet).
Here's what you're doing with spiking nets (more or less). You show them ImageNet (or any dataset) and you keep showing it to them. It will build up an internal model of what the world looks like. After training you train a second algorithm that searches which neuron encodes what. So the idea is that one of the neurons in the network will encode "I saw the letter A", another will encode "I saw B", third will encode "I saw C", so you look which neuron does that.
Spiking neural networks need time, which is another disadvantage. This is "simulated" time, but it nevertheless requires computation to happen to advance time. It takes spiking networks some amount of this simulated time to recognize things (just like animals/humans need time). So to have it recognize it you have to "put a picture in front of them for X time" (meaning keep triggering their inputs in the same way for a while), then wait to see if any of the identified neurons fires in the first, say, 5s after showing the picture.
Where spiking networks wipe the floor with convolutional nets is on unpredictable tasks. Suppose you were having an "open-ended" problem (like, say, a lifeform has). And the environment changes. A convolutional net that was trained will, quite simply, start giving random results. A spiking network will do something. Not necessarily the right thing, but it will try things.
Say you were building a robot that has to deliver supplies across Iraq. Convoluational nets won't adapt. Spiking models will adapt (assuming you let them, and I imagine DARPA will let them). Problem with letting them adapt, of course, is that you may lose control.
And before you say "what about morality ?", I would say that spiking nets are actually more moral. In both cases, spiking or convolutional, you don't actually know how it will respond to unpredictable stimuli. However, if a convolutional net is confronted with something it wasn't trained for, it will simply have random reactions (it's a robot, it'll send random instructions to the higher levels, meaning if it has a gun, it will extremely likely fire the gun, probably aimed at the first thing it recognizes), a spiking model will try something (which, of course, may be "kill all humans", but it might also decide to wait and see if there are hostile moves, or ...). The difference is the spiking model won't simply lose control. I would argue that spiking models will respond much more like soldiers would.
You should think of convolutional nets as classifiers. You train them to answer a yes/no question, and then they can respond. Spiking neural nets are more like puppies. You can train them to bark if they see a car, and then use that to detect cars, but you can also train them to retrieve a ball. (In practice you "read the mind" of the spiking model, and because it's stored in memory, that's easy)
Aside from the particular criticism he makes toward the IBM algorithm, it seems to me that the approach of jumping from one special chip to another abandons the advantages of a general purpose computer itself. If your algorithm has to be cast in silicon each time, tuning the algorithm would depend on the chip's lifecycle. Also, only those few who have the resources to build a chip would be able to supply algorithms narrowing the number of minds working on this.
That alternative I'd like to see is a general purpose highly parallel chip.
The one I know of is the Micron automaton chip. http://www.micron.com/about/innovations/automata-processing
Anyone know of anything similar?
Also, the whole "airplanes don't flap their wings" analogy can be taken too far. Little flying things are qualitatively different from big flying things. You'll notice that a lot of small artificial flying things are flapping, and the biggest natural fliers tend to glide a lot. There are other reasons why nature didn't evolve large fliers. (Although I'm willing to believe some large fliers may occasionally have hot gases shooting out of their back ends, I do not believe propulsion is their purpose.)
The fact that Truenorth can learn approximations is not really surprising, we know that thresholded units can approximate well[1]. They should have implemented compartmental neurons [2]
[1] http://en.wikipedia.org/wiki/Universal_approximation_theorem [2] http://en.wikipedia.org/wiki/Compartmental_modelling_of_dend...
What would be called 'good science'? Reading articles and spinning tales about what comes next? Regurgitating summaries of others' work?
It would be an experiment if they were testing a new model. This is a simulation.
[1] http://www.izhikevich.org/human_brain_simulation/Blue_Brain.... of Large-Scale Brain Models
This is a great asynchronous circuit power efficiency research, and not a neuroscience research at all.
Also, how do convolutional neural networks model time? I thought that was one of the benefits of spiking networks and STDP.
Isn't this better. LeCun is looking at this only from a machine learning perspective.
That being said, the brain is basically the only working example of true intelligence we have, so perhaps trying to emulate it isn't a bad idea.
But I have this sci-fi notion that eventually the AI community will produce some sort of intelligence that is unimaginably different from our current notion of a brain.
Last I heard, the behavior of biological neurons was so badly understood that even the behavior of the 300-neuron C. elegans worm could not be accurately simulated, even though the neuron connections have been fully determined. That was about 8 years ago, though.
But it seems like SyNAPTIC group tries to push this chip into common machine learning applications.
http://technology-report.com/2009/11/neuroscience-expert-dr-...