——
[0] https://web.stanford.edu/group/brainsinsilicon/index.html
the big idea here is that IF your base element (memristor crossbar here) is suitable for such rapidly reconfigurable bus architecture (which it seem like it is) then you can use it to synthesis a single neuron directly. which is a huge leap over the next best GPU/TPU based architecture based on instruction fetch-decode-execute model. based on what I have read few years ago you can have a 20M neurons simulated with memristors in about a cm2 die. that is human level integration density even if you totally ignore the vast difference in switching rate (100Hz vs 1+GHz).
Even when you look at CPU performance, you can often pinpoint bottlenecks right at the amount of available L1 or L2 cache. Cache locality has almost always been the limitations of performance, because to process data, you must first access and write data. So if memory is more closely available, then everything should always be fast.
Also, remember you cannot do software on a GPU, because GPUs, even with CUDA or OpenCL, are not built to run software for the simple fact that GPUs don't do error correction. OpenCL and CUDA will only help when processing data that can be parallelized, so where the result will not risk to be jeopardized if errors accumulate.
https://news.ycombinator.com/item?id=16436487
You can do software on a GPU, but you cannot have good guarantees.
But this is totally irrelevant anyway, the comparison we started with is to an analog alternative. Anything analog will have strictly worse noise and error problems. In neural networks, errors are probably not even a problem for analog implementation, so they definitely aren't a problem for GPU implementations.
But ok, I was not aware that gpu had error correction, how much additional transistor does it take?
The fundamental problem in modern computers isn't memory: it's in moving data around. The speed of light in a vacuum gives you only a few cm of distance to move information in a single clock cycle, and the actual electronic propagation inside the processors is substantially slower. In fact, the governing factor of the size of L1 cache is the time it takes to actually read a value. At the scale of supercomputers, the topology of the interconnect has major implications for the actual performance on HPC applications.
Saying that bringing memory closer is the determining factor in speed ignores the fact that the size of memory has implications in the time to access it. The innovation in CPUs has been about minimizing latency essentially by developing better heuristics in what it might be. GPUs innovate by not trying to minimize latency but instead trying to overprovision cores and rely on batched memory access (consequently, GPUs are not good at handling codes that rely on irregular memory access patterns).