Wouldn’t the memtransistor be faster though? Even if you use CUDA, you still have to compile down to assembly and probably use some kernel level module to communicate with the GPU.
——
[0] https://web.stanford.edu/group/brainsinsilicon/index.html
the big idea here is that IF your base element (memristor crossbar here) is suitable for such rapidly reconfigurable bus architecture (which it seem like it is) then you can use it to synthesis a single neuron directly. which is a huge leap over the next best GPU/TPU based architecture based on instruction fetch-decode-execute model. based on what I have read few years ago you can have a 20M neurons simulated with memristors in about a cm2 die. that is human level integration density even if you totally ignore the vast difference in switching rate (100Hz vs 1+GHz).