How does it do on CIFAR-10, or even better, ImageNet?
Interesting research, not sure it's a backprop alternative.
===
EDIT: accuracy on MNIST is not ~90%. It's ~85%.
How does it do on CIFAR-10, or even better, ImageNet?
Interesting research, not sure it's a backprop alternative.
===
EDIT: accuracy on MNIST is not ~90%. It's ~85%.
The obvious example is if it has different behaviour around local minima, it could be an altenate pathway out.
I have often wondered if doing training with radically different aproaches for the first few iterarions would avoid any method specific artifacts before the weights had time to denoise.
But the submission title calls it a "backprop alternative." It is not, at least not yet.
Their goal is to understand how distributed systems which cannot do backprop (the brain) can still do learning.
If we puke, the brain will not do general aversive learning, it will learn to avoid specifically the last thing eaten, because it instinctively knows about food poisoning.
Imprinting is absolutely fascinating. Some newborn animals will run a very simple pattern detector like looking for a red dot or something and use that to bootstrap their conception of their parent.
For fully general learning I have a hunch that it can be done using local history plus a semi-global reward scalar (global neurotransmittor levels).
For another, there are about 200k promotor regions (including non-coding) in the human genome.
A promotor region might have say 6 to 15 bits of information.
Can you compress 2025 or even 2024 era LLM intelligence into 3 megabit = ~400 kB ? I think not. I think a lot of compression is still possible, but 400 kB?
So I think we can box up the idea of "dirty secrets of the braing: not learning but hard coding". There is a lot of hard coding in biology, but brains are evolved specifically to enable learning within the individual lifetime instead of only learning by natural selection.
I also don't buy the following argument:
> If we puke, the brain will not do general aversive learning, it will learn to avoid specifically the last thing eaten, because it instinctively knows about food poisoning.
Each time it happens that I end up puking, I do feel aversion and try to avoid puking at all, sometimes I succeed but sometimes is just puke. There must be fundamental puke reflexes (which one fails to avoid) and avertable puke reflexes.
There are a few extra levels of interpretation (like protein synthesis) that are more like a transpiler than compression (imo), over a 4-base language that is read in a sliding window and is affected by surrounding conditions, so the same "token" sequence may produce different things depending on external factors. Some biologists I used to collaborate with talked about 7 layers to this process, I have only described one level here
> There are a few extra levels of interpretation (like protein synthesis) that are more like a transpiler than compression (imo), over a 4-base language that is read in a sliding window and is affected by surrounding conditions, so the same "token" sequence may produce different things depending on external factors. Some biologists I used to collaborate with talked about 7 layers to this process, I have only described one level here
so the same "token" sequence may produce different things depending on external factors.
yes, non-hereditary learning depends on external factors, thank you for paraphrasing me while shifting attention.
the multi-scale nature (transpilers etc.) doesn't change the theorems in probability and information theory which seriously constrain the maximum amount of information a message can store.
The whole point of a brain is that it is an organ dedicated to storing, retrieving and timely utilising information one can't afford to store in a genome.
I was never paraphrasing you, but thank you for attempting rhetorical antics?
I'm not convinced my molecules follow the laws of thermodynamics.
1. A determinist transformation can't create more information than its inputs
you can't derive (2)
2. Later stages don't contain any new information.
without restraining your model to a monotonic chain of transformations that can't tap into the information content of the environment.
We instinctively know that there are other intelligent beings, and we have the the machinery to model them. We are born with the capacity for language. It must still be learned, the specific words aren't hardcoded, but the concept of language is. We are born with the capacity to store and replay memories. The brain knows some aspects of how the world is supposed to look like visually, and if it doesn't it will try to correct that. People who used optics to see the world upside down have found that after a brief time their brain learned to flip the world right side up.
I think my initial assessment is correct: interesting research, but not an alternative, at least not yet.
I didn't see ImageNet. TinyImageNet is something else.
Training a Predictive Coding Network on ImageNet using Equilibrium Propagation Tugdual Kerjan, Rasmus Høier, Benjamin Scellier https://arxiv.org/abs/2606.03584
It's quite an undertaking.