Self-Compressing Neural Networks
arxiv.org
arxiv.org
I'd want a model that scales up and down depending on the task given at inference, and a model that doesn't have a fixed size when starting the training. Shouldn't it specialize over training progress, when seeing more tokens, and grow larger where needed? Without some human fixing a size beforehand?
Self-organization is a fascinating topic to me. This last year I've been working on Self-Organizing Gaussian Splats [0]. With a lot of squinting, this lives in a similar space as the Self-Compressing Neural Networks from the link above. The idea of the Gaussians was to build on Self-Organizing Maps (lovely 90s concept, look for some GIFs if you don't know it), and use that to represent 3D scenes in a memory-efficient way. By mapping attributes into a locally smooth 2D grid. It's quite a simple algorithm, but works really well, and better than many quite complicated coding schemes. So this has me excited that we'll (re-)discover great methods in this space in the near future.
[0]: https://fraunhoferhhi.github.io/Self-Organizing-Gaussians/
And I'd be curious of the utility of model that scales up and down at inference - if this was the case you'd still need to have storage that is the same as the maximum model size. This would essentially be useless for embedded applications, etc., unless you have heavy quantization - but quantization in a small parameter space would probably make the smaller modes useless. I could see the benefit here in terms of optimizing latency for different applications but maybe you have other ideas.
Given all that, I think training for smaller number of parameters, as noted in OP, would kind of beat out some model that scales at inference time - especially when most people know what kind of application they are aiming to build and the required level of performance.
This paper is a long way from implementing synaptic pruning/strengthening/weakening, neurogenesis, or synaptogenesis but it’s the first one I’ve seen where the network is self optimizing.
As PAC learning with autograd and perceptrons is just compression, or set shattering, this paper is more of an optimization method that reduces ANN expressiveness through additional compression. Being able to control loss of precision is exciting though.
It may help in some cases, especially for practical use cases, but their unaddressed mention of potential problems with noisy loss functions needs to be addressed.
Human biological neurons can do XOR in the dendrites without hitting the soma at all is another example.
If you haven't heard about dendritic compartmentalization and plasticity, here is a paper.
https://www.cell.com/neuron/fulltext/S0896-6273(11)00993-7
> In conclusion our results support the view that experience can drive clustered synaptic enhancement onto neuronal dendritic subcompartments, providing fundamental architecture to circuit development and function
But does it? It’s been my hypothesis for a while that every grad-trained NN is hauling around a lot of “nascent” nodes — nodes that were on their way to being useful, but haven’t received enough input yet to actually have their outputs be distinguishable from noise / ever influence the output. Sort of the neuroplastic equivalent of an evolutionary pre-adaptation.
If such nodes exist in NNs, they would be important to decreasing training time to learning new concepts given further training; but if there will be no more training, then they could be pruned for literally no change in expressivity (i.e. the optimality of the NN as an autoencoder of the existing training data.)
Consider when you use 'partial connectivity', E.G. convolution or pooling layers for local feature extraction on say MNIST.
While useful, those partial connection layers are explicitly used because fully connected layers do not have translational invariance.
So with a fully connected network, shifting the letter 'i' a few pixels to the right wouldn't match.
We choose to discard some of those connections for local feature detection. But as the reason that the fully connected model lacks translational invariance is because it maintains that position data.
Note how that is more 'expressive', even if counterproductive for the actual use case.
Another lens is the fact that neural networks have extreme simplicity bias. In that they learn only the simplest features to solve a task at hand.
If you want to recognize an i, irrespective of the translational location, that bias is useful. But you 'throw away' (in a very loose sense) the positional data to do so.
Horses for courses, not good vs bad.
"Naud and Sprekeler (2018) suggest that this could be achieved using a synaptic strategy that facilitates summation for simple action potentials arriving on the basal dendrites and depresses faster burst-like events arriving on the distal tuft"
Oh, its frequency multiplexing with a band pass filter. Same trick the analog phone system used to reduce the amount of wire needed in the network. Same problem, same solution. Convergent evolution.
I wonder if there's ways to do phreaking on neurons.
https://www.sciencedirect.com/science/article/pii/S030645222...
Similarly, do you think that spiking networks are important, or just a specific mechanism used in the brain to transmit information, which dense (or sparse) vectors of floats do in artificial neural networks?
If your goal is to produce a useful model on real hardware and it works...no
Remember the constraints of ANNs being universal approximaters (in theory)
1) The function you are learning needs to be continuous 2) Your model is over a closed, bounded subset of R^n 3) The activation function is bounded and monodial
Obviously that is the theoretical UAT constraints. For gradient decent typically used in real ML models, the constraint of finding only smooth approximations of continuous functions can be problematic depending on your needs.
But people leveraged phlogiston theory for beer brewing with great success and obviously Newtonian Mechanics is good enough for many tasks.
SNNs in theory should be able to solve problems that are challenging for perceptron models, but as I said, features like riddled basins are problematic so far.
Seems like a bad limitation when you try to model reasoning based on facts and logic, there are many things there that are just true or false and no spectrum to it. There is no "kinda true" in those circumstances, you should only get 1 or 0 and never any value between.
While not practical to find or use, any feed forward network supervised is effectively a paramedic linear regression.
Think of an Excel line graph, drawing lines between points, with the above the line being 'true', or when the soma fires.
That is how perceptrons work.
Single layer perceptrons cannot represent linearly inseparable functions like XOR or band pass.
A single biological neurons can use the timing of pulses, band pass, change the rate of pulses etc... before it ever reaches the soma.
Not all problem can be reduced to decision problems and not all of them can use constant depth threshold circuits, which hard attention is.
An LLM can be a very reliable threshold or majority gates as an example, but cannot generalize PARITY.
Basically statistical learning inherited the same limits of statistics.
"This statement is 'False'" is a good paradox to use as a lens.
Eg dropout was (allegedly) inspired by our doubled up chromosomes and evolutionary selection.
We need both to work on improving what we have that works, and to explore other avenues and inspirations (both to try entirely new things, and to improve the things we already have working in new ways). I don’t think it wise to throw out what we have working to try again with something biologically inspired, but I also don’t think it wise to say ok we’ve learned enough from biology, let’s focus purely on what we have now, when we don’t understand so much about biological brains, intelligence, and consciousness.
That’s not what I said.
Try different things, choose what works, as opposed to trying to imitate biology for the sake of imitating biology.
Right now, we know that biology works, because of animal and human intelligence. We don’t yet know if our other approaches have the ability to eventually lead to that.
[0] https://en.wikipedia.org/wiki/File:BirdVisualPigmentAbsorban... [1] https://en.wikipedia.org/wiki/File:Cones_SMJ2_E.svg
Isn't this already accomplished via weights?
I don't think that's actually a good goal. I suspect the whole term 'neural network' is just misleading and leads to these kinds of misconceptions.
'Neural networks' are mostly just matrix multiplications interleaved with some simple non-linear functions like \x -> max(0, x). Nothing biological about that.
https://arxiv.org/abs/1712.01312
Abstract: "We propose a practical method for L0 norm regularization for neural networks: pruning the network during training by encouraging weights to become exactly zero. Such regularization is interesting since (1) it can greatly speed up training and inference, and (2) it can improve generalization. [...]"
My take on it: I find it difficult to generalize the notion of layer removal when the bit depth of that layer goes to zero. It's wouldn't be straight forward although the authors provide equation 5. It feels like lot of information is missing in this work to even reproduce it. And authors do only 1 case study.
I believe some implementation is required to understand the authors completely. Example, optimizer modification for layer when it is removed in training.
https://x.com/realGeorgeHotz/status/1819963680739512550
> This is one of the coolest papers I've seen in a while. "Self-Compressing Neural Networks" is dynamic quantization-aware training that puts size (in bytes) of the model in the loss! > My implementation (in @__tinygrad__):
https://github.com/geohot/ai-notebooks/blob/master/mnist_sel...
If i can sort the lines of the matrix, which is probably defined by how the token embedding is setup, i could potentially zero out weights which do not contribute at all and have areas of zeros i could mark and skip?
Edit: In general the more a compression function understands of what your goals are the better, so it is naturally advantageous to make the compression function look like training since then it is fully aware of what to optimize for.
My meek opinion is this is obvious. Human-level intelligence requires at most 20 watts and substrate no more complicated than can be constructed from simple organic molecules in a dirty environment.
What is possible with 20 kilowatts and wafer fabricators?
Anyway, OP is about lossy compression. I can't fully follow it but they talk about techniques for mitigating loss later in the paper.