IBM Chip Processes Data Similar to the Way Your Brain Does
technologyreview.com
technologyreview.com
1. Building special-purpose hardware for neural nets is a good idea and potentially very useful.
2. The architecture implement by this IBM chip, spike-and-fire, is not the architecture used by the state-of-the-art convolutional networks, engineered by Alex Krizhevsky and others, that have recently been smashing computer vision benchmarks. Those networks allow for neuron outputs to assume continuous values, not just binary on-or-off.
3. It would be possible, though more expensive, to implement a state-of-the-art convnet in hardware similar to what IBM has done here.
Of course, just because no one has shown state-of-the-art results with spike-and-fire neurons doesn't mean that it's impossible! Real biological neurons are spike-and-fire, though this doesn't mean the behavior of a computational spike-and-fire 'neuron' is a reasonable approximation to that of a biological neuron. And even if spike-and-fire networks are definitely worse, maybe there are applications in which the power/budget/required accuracy tradeoffs favor a hardware spike-and-fire network over a continuous convnet. But it would be nice for IBM to provide benchmarks of their system on standard vision tasks, e.g., ImageNet, to clarify what those tradeoffs are.
I am doing some experiments in this area, and would encourage anyone thinking of doing hardware to look at this aspect before investing the R&D to do hardware! If this knowledge can really be compressed it could be a massive reduction in complexity to implement in hardware...
I am a bit biased on this topic (finishing a talk about this exact topic for EuroScipy now) but I find the connections interesting at least.
http://www.research.ibm.com/software/IBMResearch/multimedia/...
At the time, functional programming was not exactly mainstream and many of the concurrency concepts we take for granted today from web programming were just research. So of course nobody listened to ranters like me and the world plowed its resources into GPUs and other limited use cases.
My take is that artificial general intelligence (AGI) has always been a hardware problem (which really means a cost problem) because the enormous wastefulness of chips today can’t be overcome with more-of-the-same thinking. Somewhere we forgot that, no, it doesn’t take a billion transistors to make an ALU, and no matter how many billion more you add, it’s just not going to go any faster. Why are we doing this to ourselves when we have SO much chip area available now and could scale performance linearly with cost? A picture is worth a thousand words:
http://www.extremetech.com/wp-content/uploads/2014/08/IBM_Sy...
I can understand how skeptics might think this will be difficult to program etc, but what these new designs are really offering is reprogrammable hardware. Sure, we only have ideas now about what network topologies could saturate a chip like this, but just watch, very soon we’ll see some wizbang stuff that throws the network out altogether and uses content addressable storage or some other hash-based scheme so we can get back to thinking about data, relationships and transformations.
What’s really exciting to me is that this chip will eventually become a coprocessor and networks of these will be connected very cheaply, each specializing in what are often thought of as difficult tasks. Computers are about to become orders of magnitude smarter because we can begin throwing big dumb programs at them like genetic algorithms and study the way that solutions evolve. Whole swaths of computer science have been ignored simply due to their inefficiencies, but soon that just won’t matter anymore.
It's really a combination of memory latency and pipelining.
Memory latency is absolutely terrible compared to processor speed, and that has nothing to do with Moore's law. It's 60ns to access main memory, which is ballpark 150 cycles. If you have no caches, your 2.5Ghz processor is basically throttled to 16Mhz. You can buy some back with high memory bandwidth and a buffer (read many instructions at a time). But if you have no predictor, every taken branch flushes the buffer and costs an extra 150 cycles- in heavily branched code your performance approaches 8Mhz.
Then think about pipelining. We don't pipeline because Moore's law has ended. We pipeline because a two-stage pipeline is 200% as fast as an otherwise identical unpipleined chip. A sixteen-stage pipeline is 1600% as fast. Why the hell wouldn't you pipeline? Now, of course in the real world branched code can tank a deep pipeline. Which is where the branch predictor comes in, buying back performance.
http://stackoverflow.com/questions/4087280/approximate-cost-...
No. This is only true if every instruction tries to access memory.
>>> We pipeline because a two-stage pipeline is 200% as fast as an otherwise identical unpipleined chip. A sixteen-stage pipeline is 1600% as fast.
No. First of all, each stage in the pipeline will be equal to the slowest stage. Second, there will be significant overhead of passing data through pipeline registers, and of control logic for those registers.
The reason we saw 32 stage pipelines in P4 was mostly marketing: "megaherz race" between AMD and Intel.
But you can be certain that AMD and Intel do not design 20+ stage pipelines for some measly 10% performance uplift. The overhead of the pipeline infrastructure is nowhere near the performance gain. Consider Haswell has an IPC around 2 instructions per cycle. With a ~20 stage pipeline, they are indeed far outstripping the performance of "Haswell minus pipelining".
As for the super-deep pipeline in the P4, the consensus I hear is that Intel expected frequency to keep scaling, and as such the P4 was a future-looking architecture designed to scale to 10GHz and beyond.
Every instruction must be loaded from memory in order to execute it. Hence instruction caches.
This is really exciting stuff, I can't help but think a marriage of this approach with HP's memristor technology would bring us screaming along an amazing architecture path for the next several decades.
But then again, I'm concerned that the limited use cases for this being presented are basically already performed by various custom (and cheap and power efficient) DSPs. Is all that's really being envisioned here just a lower power alternative to DSPs? I think the vision can be much bolder.
Imagine if you were a reference librarian, asked for facts like some kind of ancient Google. Suppose your library was the size of your bedroom- you could very quickly find facts. You only have to cross the room. Now suppose you are right in the middle of the Library of Congress. You are smack dab in the middle of it- you are where the memory is. But you're still going to spend half your time just running about the building due to its sheer size!
The only ways to solve that problem are:
- Make memory smaller. Engineers have been hard at work at this for decades.
- Use less memory. This is slower.
- Use a memory hierarchy. This is what we do today, and is analogous to you sitting in a bedroom-sized library with the Library of Congress just down the street, and a young courier who fetches you books from it.
The other challenge is speed. We can't have a huge pile of registers because fast memory is less-dense than slow memory. So 1KB of CPU registers occupies a lot more space than 1KB of DRAM- but DRAM is a poor choice for registers because of how slow it is.
- use more CPUs
With a billion perfectly cooperating (that's the research problem) librarians, searching the library of congress is way faster.
Multicore performance doesn't scale linearly because 1) adding more cores has rapidly diminishing returns on performance for most problems (http://en.wikipedia.org/wiki/Amdahl's_law) and 2) the cost of coherency is exponential with the number of cores.
The philosophy of today's GPU architecture is basically quite simple: Maximize memory throughput by using the fastest RAM that's still cheap enough for consumers, then maximize die space for the ALUs by letting bundles of them share scheduler, register blocks and cache. I was first very skeptical about this too, but to my experience it has proven quite effective - even parallel algorithms that are not ideal for this architecture still profit from the raw power, and they continue getting benefits when you buy new cards, in a fashion that's much closer to Moore's law than CPUs develop.
The architecture certainly isn't ideal and would be solved by an architecture like in your link (to which Parallela also comes quite close btw), and I can well imagine that this is where we're heading given another 5-10 years (see Parallela, to some extent Knight's Landing). However it's also feasible that the GPU's ALU maximisation game will win out, especially once 3D-stacked memory comes into play.
Since 2008 there have been many papers about NNs implemented on GPU and I'd love to know what's the current status there, especially compared to the very powerful Power8 architecture.
Meaning I have no idea how this signals the beginning of a new era of more intelligent computers as the chip provides nothing to advance the state of the art on this front. Unless I am missing something?
Of course, there could be some new and interesting uses in embedded devices where sheer throughput doesn't matter so much as total power usage, for moderate processing power. For example, the AI in Roomba and similar robot vacuums is pretty rudimentary, so appliances like that could maybe get a boost from this.
Just an uneducated wild-thought.
Neat developments, excited to see how they shake out.
See this paper at ASPLOS '14 for details:
http://hips.seas.harvard.edu/content/asc-automatically-scala...
A recent discovery by Leon Chua has shown that synapses and neurons can be directly replicated using Memristors [1]. Memristors are passive devices which may be much simpler to build in the scale of neurons compared to transistors.
The comparison of this chip's performance with that of a nearby traditionally-chipped laptop is questionable. A couple of paragraphs later it says that the chip is programmed using a simulator that runs on a traditional PC. So I'm guessing the 100x slowdown is because the traditional PC is simulating the neural-net hardware, rather than using optimized software of its own.
Yes, this is important research, but engineer-speak piped through hype journalists will always paint an entirely unrealistic and overoptimistic picture of what's really going on.
Are you sure about that? CPU speeds have not improved in years. We appear to have hit a maximum, at least for now. (Of course I can't predict the future, but it's been years now and no change.)
Otherwise we would still be using (very cheap) Pentiums IV.
In a way, it's a testament to human ingenuity that CPUs have kept improving they way they have when the brute force way of increasing performance was not as viable as before.
Personally I believe digital computation has only a niche applicability in the limit, the degrees of freedom from analog processing are just so much higher, even in the presence of noise.
Indeed, I could even claim that given how little we know, the actual "real processing" happen in the brain wind-up being much less than it seems. But yes, it appears that whatever the brain does is fantabulously more complex than any chip that's even being sketched today.
What this looks like is a chip that does some canned machine learning routines. It seems sad to have to hype a parallel chip of this sort this way. But it would be sad if the chip itself is hard corded for just whatever fake-brain computations its creators thought were right (I've scanned several pages deep for real information on the chip but it comes back hype and more hype). The thing is it's actually possible to build a more general kind of parallel chip - a cellular automaton on chip such as Micro is doing, see: http://www.micron.com/about/innovations/automata-processing.
Also, the Wikipedia page gives the impression this is mostly an exercise in seeing if they can scale chip to neural scale. http://en.wikipedia.org/wiki/SyNAPSE
Thus, while the topics are fascinating and these appear to be impressive strides, journalists need to be careful not to hyperbolize.
I was on a DARPA neural network tools advisory panel for a year in the 1980s, developed two commercial neural network products, and used them in several interesting applications. I more or less left the field in the 1990s but I did take Hinton's Coursera class two years ago and it is fun to keep up.
Well I found this little sound-byte from a link in the original article. I didn't find it particularly original.
From [0] > “Programs” are written using special blueprints called corelets. Each corelet specifies the basic functioning of a network of neurosynaptic cores. Individual corelets can be linked into more and more complex structures—nested, Modha says, “like Russian dolls.”
The term 'Russian doll' evoked recursion (and distant memories of my late grandfather), very common even 50 odd years ago.
[0] http://www.technologyreview.com/news/517876/ibm-scientists-s...
Interesting, I did not know that we already know how the brain 'processes data'.
http://dx.doi.org/10.1126/science.1254642
More broadly, I don't understand why HN seems to prefer press pieces (so often containing more inaccuracies than useful information) to the papers on which they're based.
In this case, even if you can't access the full text, the single-paragraph abstract contains all of the new information in the 12-paragraph Tech Review story.