'Memtransistor' Forms Foundational Circuit Element to Neuromorphic Computing
spectrum.ieee.org
spectrum.ieee.org
Now we have something called a "memtransistor". We also have something called a "memristor".
But invariably in these discussions, there's another device that exists that almost never seems to get a mention. It's called a "memistor":
https://en.wikipedia.org/wiki/Memistor
It was first developed in 1960, and used for a couple of related hardware neural network architectures - ADALINE and MADALINE.
What is also interesting about this device, is that it can be relatively easily built by a hobbyist, as shown by:
http://www-isl.stanford.edu/~widrow/papers/t1960anadaptive.p...
Strangely, though - the memistor is hardly ever mentioned; it does have some downsides (mainly difficulty to miniaturize), so maybe that's it? Maybe somebody here can play with it...
Interestingly, on a side note - it's possible to homebrew a memristor as well:
The big deal about memristors for AI is that they have memory and therefore produce different readout outputs for different input sequences over time -- out-of-the-box, without requiring any training -- and in a way that can make sequences linearly separable under fairly general conditions. For instance, in the Nature paper I linked to above, the researchers took a network of memristors, which they call the "reservoir," added a linear layer with a SoftMax on top of the reservoir, trained this hybrid network on a lower-resolution variant of MNIST (feeding pixel values over time, as varying voltages), and achieved classification accuracy superior to a tiny neural net despite having only 1/90th the number of neurons.[1] Note that they only trained the added layer; they did not have to train the reservoir. Figure (a) in this image has a simplified diagram of the reservoir + added layer architecture: https://www.nature.com/articles/s41467-017-02337-y/figures/1 -- only the matrix Θ had to be learned. The potential, over time, is for having highly scalable hardware neural-net components for learning to recognize and work with sequences.
PS. For clarity's sake, I'm ignoring a lot of important details and playing fast and loose with language. If you're really curious about this, please read the IEEE article and the Nature paper.
[1] https://news.engin.umich.edu/2017/12/new-quick-learning-neur...
As you can see for yourself in figure [1]c, the memristor net used in that paper has only 15 units (yes, fifteen): five memristors in the "reservoir" and 10 output neurons. The entire net is one linear transformation followed by a SoftMax, with a total of 5 × 10 = 50 parameters.
It's an impressive result that only hints at the things that will be possible when we have nets with millions or billions of memristor units.
[1] https://www.nature.com/articles/s41467-017-02337-y/figures/1
——
[0] https://web.stanford.edu/group/brainsinsilicon/index.html
the big idea here is that IF your base element (memristor crossbar here) is suitable for such rapidly reconfigurable bus architecture (which it seem like it is) then you can use it to synthesis a single neuron directly. which is a huge leap over the next best GPU/TPU based architecture based on instruction fetch-decode-execute model. based on what I have read few years ago you can have a 20M neurons simulated with memristors in about a cm2 die. that is human level integration density even if you totally ignore the vast difference in switching rate (100Hz vs 1+GHz).
Even when you look at CPU performance, you can often pinpoint bottlenecks right at the amount of available L1 or L2 cache. Cache locality has almost always been the limitations of performance, because to process data, you must first access and write data. So if memory is more closely available, then everything should always be fast.
Also, remember you cannot do software on a GPU, because GPUs, even with CUDA or OpenCL, are not built to run software for the simple fact that GPUs don't do error correction. OpenCL and CUDA will only help when processing data that can be parallelized, so where the result will not risk to be jeopardized if errors accumulate.
https://news.ycombinator.com/item?id=16436487
You can do software on a GPU, but you cannot have good guarantees.
But this is totally irrelevant anyway, the comparison we started with is to an analog alternative. Anything analog will have strictly worse noise and error problems. In neural networks, errors are probably not even a problem for analog implementation, so they definitely aren't a problem for GPU implementations.
But ok, I was not aware that gpu had error correction, how much additional transistor does it take?
The fundamental problem in modern computers isn't memory: it's in moving data around. The speed of light in a vacuum gives you only a few cm of distance to move information in a single clock cycle, and the actual electronic propagation inside the processors is substantially slower. In fact, the governing factor of the size of L1 cache is the time it takes to actually read a value. At the scale of supercomputers, the topology of the interconnect has major implications for the actual performance on HPC applications.
Saying that bringing memory closer is the determining factor in speed ignores the fact that the size of memory has implications in the time to access it. The innovation in CPUs has been about minimizing latency essentially by developing better heuristics in what it might be. GPUs innovate by not trying to minimize latency but instead trying to overprovision cores and rely on batched memory access (consequently, GPUs are not good at handling codes that rely on irregular memory access patterns).
To purchase: http://www.bioinspired.net/products-1.html