Training of Physical Neural Networks
arxiv.org
arxiv.org
The very thing that makes it so powerful and efficient is also the thing that make it uncopiable, because sensitivity to tiny physical differences in the devices inevitably gets encoded into the model during training.
It seems intuitive this is an unavoidable, fundamental problem. Maybe that scares away big tech, but I quite like the idea of having invaluable, non-transferable, irreplaceable little devices. Not so easily deprecated by technological advances, flying in the face of consumerism, getting better with age, making people want to hold onto things.
And a more approachable article: https://www.damninteresting.com/on-the-origin-of-circuits/
Another idea is to just train a whole bunch of them individually, like putting your chips in school. :-D
We are not even born with what you might consider basic mental faculties, for example it might seem absurd, but we have to learn to see... We are born with the "hardware" for it, a visual cortex, an eye, all defined by our genes, but it's actually trained from birth, there is even a feedback loop that causes the retina to physically develop properly.
Likewise, kittens with an eye patch over an eye in the same time period remain blind in that eye forever.
Children who were "raised in the wild" or locked in a room by themselves have shown to be incapable of learning full human language.
The working theory is that our brains can only learn certain skills at certain times of brain development/ages.
I suspect there may be trade off undergoing evolutionary selection here, where for some organisms a behaviour is more important from the offset, it's worth encoding more of the behaviour into genes, at what cost I wonder?
It's also possible there is some other mechanism going on at an embryonic stage, a kind of pre-training.
I suspect some of the division is also defined by how complex the task is, or how sensitive the model is to it's own neurons (kind of like PNN). I don't have a well rounded argument, but my instinct is that encoding or pre-training walking is far easier than seeing. Not to mention basic quadrupedal walking/standing is far easier than bipedal, they can learn the more complex coordinated movements after.
A ton of that is probably encoded elsewhere, but no doubt the brain plays a huge part. And somehow, it's all reconstructed for each new "device".
There is a great write up of this in this old blog post: https://www.damninteresting.com/on-the-origin-of-circuits/
With ANN you can do it one time and then clone the result for negligible energy cost.
Maybe training a batch of PNNs in parallel could save some of the energy cost, but I don't know how feasible that is considering they could behave slightly differently during training causing divergence... Now that sarcastic comment at the bottom of this thread is starting to sound relevant "Schools".
That entirely depends on how many inferences the model will perform during its lifecycle. You can find different estimates for the energy consumption of ChatGPT, but they range from something like 500-1000 MWh a day. Assuming an electricity price of $0.165 per kWh, that would put you at roughly $80,000 to a $160,000 a day.
Even at the lower end of $80,000 a day, you'll reach your $4 Million in just 50 days.
With PNN you would have to multiply n by 1-4 million, training cost explodes.
Having to do that in each instance is still really cumbersome for cheap mass deployment compared to just making a digital-style exact copy, but then again I guess a main argument for wanting these systems is that they'd be doing things unachievable in practice on digital computers.
In some cases one might be able to distill to digital arithmetic after the heavy parts of the optimization are done, for replication, distribution, better access for software analysis, etc.
I think eventually we'll get to the point where we do a stage of pretraining on noisy digital hardware to create a transferrable network, then fine tune it on the analog system.
[1]: https://github.com/neuromorphs/nir (disclaimer: I'm one of the authors)
I am trying to understand what format does a node take in PNNs. Is it a transistor? Or is it more complex than that? Or, is it a combination of a few things such as analog signal and some other sensors which work together to form a single node that looks like the one we are all familiar with?
Can anyone please help me understand what exactly is "physical" about PNNs?
Last year, researchers from the University of Sydney and UCLA used NWNs to demonstrate online learning of handwritten digits with an accuracy of 93%.
This is a pretty big problem, though if you use information-bottleneck training you can train each layer simultaneously.
Why this isn't called Hardware Neural Nets is beyond me.