This paper is a long way from implementing synaptic pruning/strengthening/weakening, neurogenesis, or synaptogenesis but it’s the first one I’ve seen where the network is self optimizing.
This paper is a long way from implementing synaptic pruning/strengthening/weakening, neurogenesis, or synaptogenesis but it’s the first one I’ve seen where the network is self optimizing.
As PAC learning with autograd and perceptrons is just compression, or set shattering, this paper is more of an optimization method that reduces ANN expressiveness through additional compression. Being able to control loss of precision is exciting though.
It may help in some cases, especially for practical use cases, but their unaddressed mention of potential problems with noisy loss functions needs to be addressed.
Human biological neurons can do XOR in the dendrites without hitting the soma at all is another example.
If you haven't heard about dendritic compartmentalization and plasticity, here is a paper.
https://www.cell.com/neuron/fulltext/S0896-6273(11)00993-7
> In conclusion our results support the view that experience can drive clustered synaptic enhancement onto neuronal dendritic subcompartments, providing fundamental architecture to circuit development and function
But does it? It’s been my hypothesis for a while that every grad-trained NN is hauling around a lot of “nascent” nodes — nodes that were on their way to being useful, but haven’t received enough input yet to actually have their outputs be distinguishable from noise / ever influence the output. Sort of the neuroplastic equivalent of an evolutionary pre-adaptation.
If such nodes exist in NNs, they would be important to decreasing training time to learning new concepts given further training; but if there will be no more training, then they could be pruned for literally no change in expressivity (i.e. the optimality of the NN as an autoencoder of the existing training data.)
Consider when you use 'partial connectivity', E.G. convolution or pooling layers for local feature extraction on say MNIST.
While useful, those partial connection layers are explicitly used because fully connected layers do not have translational invariance.
So with a fully connected network, shifting the letter 'i' a few pixels to the right wouldn't match.
We choose to discard some of those connections for local feature detection. But as the reason that the fully connected model lacks translational invariance is because it maintains that position data.
Note how that is more 'expressive', even if counterproductive for the actual use case.
Another lens is the fact that neural networks have extreme simplicity bias. In that they learn only the simplest features to solve a task at hand.
If you want to recognize an i, irrespective of the translational location, that bias is useful. But you 'throw away' (in a very loose sense) the positional data to do so.
Horses for courses, not good vs bad.
"Naud and Sprekeler (2018) suggest that this could be achieved using a synaptic strategy that facilitates summation for simple action potentials arriving on the basal dendrites and depresses faster burst-like events arriving on the distal tuft"
Oh, its frequency multiplexing with a band pass filter. Same trick the analog phone system used to reduce the amount of wire needed in the network. Same problem, same solution. Convergent evolution.
I wonder if there's ways to do phreaking on neurons.
https://www.sciencedirect.com/science/article/pii/S030645222...
Similarly, do you think that spiking networks are important, or just a specific mechanism used in the brain to transmit information, which dense (or sparse) vectors of floats do in artificial neural networks?
If your goal is to produce a useful model on real hardware and it works...no
Remember the constraints of ANNs being universal approximaters (in theory)
1) The function you are learning needs to be continuous 2) Your model is over a closed, bounded subset of R^n 3) The activation function is bounded and monodial
Obviously that is the theoretical UAT constraints. For gradient decent typically used in real ML models, the constraint of finding only smooth approximations of continuous functions can be problematic depending on your needs.
But people leveraged phlogiston theory for beer brewing with great success and obviously Newtonian Mechanics is good enough for many tasks.
SNNs in theory should be able to solve problems that are challenging for perceptron models, but as I said, features like riddled basins are problematic so far.
Seems like a bad limitation when you try to model reasoning based on facts and logic, there are many things there that are just true or false and no spectrum to it. There is no "kinda true" in those circumstances, you should only get 1 or 0 and never any value between.
While not practical to find or use, any feed forward network supervised is effectively a paramedic linear regression.
Think of an Excel line graph, drawing lines between points, with the above the line being 'true', or when the soma fires.
That is how perceptrons work.
Single layer perceptrons cannot represent linearly inseparable functions like XOR or band pass.
A single biological neurons can use the timing of pulses, band pass, change the rate of pulses etc... before it ever reaches the soma.
Not all problem can be reduced to decision problems and not all of them can use constant depth threshold circuits, which hard attention is.
An LLM can be a very reliable threshold or majority gates as an example, but cannot generalize PARITY.
Basically statistical learning inherited the same limits of statistics.
"This statement is 'False'" is a good paradox to use as a lens.
I don't think that's actually a good goal. I suspect the whole term 'neural network' is just misleading and leads to these kinds of misconceptions.
'Neural networks' are mostly just matrix multiplications interleaved with some simple non-linear functions like \x -> max(0, x). Nothing biological about that.
Eg dropout was (allegedly) inspired by our doubled up chromosomes and evolutionary selection.
We need both to work on improving what we have that works, and to explore other avenues and inspirations (both to try entirely new things, and to improve the things we already have working in new ways). I don’t think it wise to throw out what we have working to try again with something biologically inspired, but I also don’t think it wise to say ok we’ve learned enough from biology, let’s focus purely on what we have now, when we don’t understand so much about biological brains, intelligence, and consciousness.
That’s not what I said.
Try different things, choose what works, as opposed to trying to imitate biology for the sake of imitating biology.
Right now, we know that biology works, because of animal and human intelligence. We don’t yet know if our other approaches have the ability to eventually lead to that.
[0] https://en.wikipedia.org/wiki/File:BirdVisualPigmentAbsorban... [1] https://en.wikipedia.org/wiki/File:Cones_SMJ2_E.svg
Isn't this already accomplished via weights?