The Scale of the Brain vs. Machine Learning (2022)
beren.io
beren.io
Anyone who has ever had kids knows this. Little kids learn language and fully developed world models on a tiny amount of data compared to the massive hoards we shovel into our training systems. That data is also noisy and only loosely curated.
The author does mention this but I think he actually doesn’t give it enough importance. A person could read for their entire life and not approach what we shovel into GPT-3.5 or llama2, yet a three year old builds a trained language model from what must be near scratch.
It must be near scratch for the initial condition because the human genome is nowhere near large enough to contain that much “priming.” There can’t be a “base model” analogue or anything like that in there.
Yeah and the brain only needs 20 watts of power.
If he/she learns it too late, they’ll never master it. Perfect pitch is an example, ‘perfect’ motor skills for something another. Language another.
Finetuning happens as a teenager/adult.
They don’t learn language in isolation. I doubt a human brain in a jar that was feed only literature would learn very well, if at all.
Being able to identify and discard unneeded data in realtime seems like a huge step for AI that hasn't been really implemented yet.
At no point in our evolutionary history have we had the leisure to absorb the vast quantities of energy that AI supercomputers can. I think this probably has a lot to do with why we are so effective at learning on little data; even when large quantities of data are available, it seems we may not have the energy to take it on board. We ignore and forget most of it.
Maybe if we made a real effort to develop energy efficient AI, that design limitation would help us develop AI that requires far less training data.
Yes, children can get some idea of a language, but they need a lot more experience and exposure to reach some level of proficiency in the language and this mostly includes supervised training in the school. They must read multiple books, have conversations etc to be able to really express themselves. It takes 10 years
The difference is that LLM don't try to have such iterative intelligence like humans do but are more like knowledge boxes. Perhaps it would take a lot less data to gain the same level of language understanding the 3 year olds have.
Verbal (and bodily) communication is quite different from written text and its largely artificially imposed structure. Spoken language is very ungrammatical and directly litterated conversational speech is a total mess on writing standards (and a lot of information is lost in litteration).
Even though e.g. Chomskian linguistics make the claim that human language is somehow a special innate syntactic construct (recursive language in the Chomsky hierarchy), I find the empirically stronger view is that the more formal grammars and vocabularies that sometimes are thought of as "language" are very far from how language is actually used when this structure is not imposed.
This quite rigid structure of "correct" written language is also likely a major factor to why LLMs are so good at it.
My 10 year old daughter doesn’t read as much as she should (too much gaming) but gets 100% on science and history tests and definitely would outperform GPT-4 on reasoning. She’d have no problem telling me how many sisters Sally has. :) A human brain consumes 20-40 watts too, not hundreds.
I bet total power consumption for a human brain age 0-10 is substantially less than what it takes to train a medium sized base model. That’d be an interesting bit of math to add.
GPT-4 will win on breadth of knowledge but that’s just because it’s a lossy compression blob containing a dump of the Internet. We are talking about intelligence (GI) here not just being a big dumb know it all.
Brains (and the nervous system and body in general) have many crucial differences from the (current mainstream) ANNs, beyond the Assumptions listed in TFA. Brain connections are higly recurrent at multiple levels, they are magnitudes slower than GPUs but they are dynamical in the state (e.g. temporal summation), there's a lot of information encoded outside the synaptic connections (e.g. concentration of transmitters/modulators in intercellular fluid), the synaptic connections are more complicated than just a scalar weight (modulated by transmitters etc), etc etc.
Also all the rage in contemporary views of animal intelligence is "embodiment" that stresses that there's more to animal intelligence than the brain.
Similar, quite apples to oranges, comparisons were made to von Neuman computation architectures (CPU FLOPS and hard drive/memory bits) when the "computer metaphor" of the brain/mind was the trend. In general there's a tendency to compare human intelligence to whatever is seen as the state-of-the-art in machines (e.g. clockworks and telephone networks back in their respective days). These probably give some intuition, but shouldn't be taken too seriously for e.g. projecting when machines and humans are at parity.
We specialize but we all have similar brains.
The last point about IQ at the very end of the article is a thought I’ve had before and is a reason I remain skeptical of the importance assigned to IQ by many. IQ seems to be a real measure of something but the variance both within and between people and the way that variance behaves is sus. I think there’s something more complex going on with IQ, probably more than one thing.
Conceptualized differently the “simulator“ in your brain that hallucinates futures is more accurate for smarter people than it is for let’s say less capable people. And some people don’t seem to have a simulator at all.
The more things and further into the future you can forecast the “smarter” you are.
I find this compelling because exceptionally smart people also seem to have exceptionally different ways of interpreting the future and also our more likely to hallucinate stuff that is totally false or incoherent. This is why you often see “geniuses” in the physical sciences struggle with social things.
The same as true people who are geniuses, socially often struggle, with other concepts outside of their narrow expertise, but are generally better and faster at learning or at least pretending better than other people with lower IQ at those general things.
It's not necessarily a bad thing, and those people still deserve to live happy lives, of course!
I can't work very well at all, but I still notice that many people just live by "ignorance is bliss" to the point where they almost try not to know anything. They don't want to know why anything works the way it does, they don't ever reflect on themselves or their actions, they don't ever perform much introspection or anything. It's almost like they operate purely on instinct. Is there a name for this phenomenon?
Maybe this is an education issue? Back when I was in school, they never taught me how to learn, they just wanted me to memorize things and then repeat them back. I don't know how that would explain the difference though, since plenty of neurotypicals are served "well enough" by school. Comparing my own experiences is sort of a slippery slope since I'm neurodivergent myself.
Like Usain Bolt runs a lot faster than me, but hey we both run so I would argue I am 95% of his performance level even though it takes twice as long for me to go 100m. I have not outsourced running to him.
E.g. Magnus Carlsen will beat me every time in chess, but compared to Stockfish me and Magnus are roughly on the same level.
A trait like intelligence is only valuable to the extent it increases an organisms fitness or number of nonsterile offspring. Witness millions of years animal evolution of organisms that had better or equivalent visual systems to humans but paltry brains.
There are some possible examples of data augmentation that might be interesting to consider though. For example, our eyes are constantly making small involuntary movements at a rate we mostly do not consciously notice. In that sense we're getting a feed of slightly transformed and translated images every few milliseconds
Thank you!
Some interesting examples: 1-Individual Purkinje cells, in isolation, learn to respond to input patterns. Meaning no synapses or other cells, the learning is happening within that cell.
2-Until relatively recently, the equipment to detect electrical activity wasn't sensitive enough so it wasn't known that dendrites have localized spiking. The collection of inputs to the dendrites are getting pre-processed either linearly or non-linearly (depending on region/function/type of cell/etc.) before being forwarded to the neurons nucleus.
3-Astrocytes have intra and inter calcium wave signaling. The intra cell signaling is localized within "proceses" (small extensions and/or compartments of the larger cell). This signaling has now been linked to sensory information processing as well as learning, meaning it's computational and separate from neuronal computation. The equipment to directly capture this activity doesn't exist so understanding is limited (studies have been based on knocking out certain capabilities and seeing the result vs direct sampling of the waves).
My opinion on comparisons: 1-A single cell in the brain is performing many more valuable computations than a single artificial neuron.
2-The brain probably also has many excess cells so a loss of a single cell doesn't have an impact. Which partially negates #1.
3-The way the brain process information could be more effective than an ANN for some functions, and it could be less effective for other functions. I think there are too many unknowns in this area to try to line it all up and compare.
> However, current visual recognition and image generation is already really good, and seems closer to parity with humans
I doubt that visual recognition is close to parity with humans. I think someone would have to define precisely what is included in "visual recognition" and what is not. It seems that "recognition" for humans means more (resolves more info about the object/context/environment) than what it means to an ANN.
I'm sorry. What? If you're starting like this I will stop reading.
By definition we are not AGI. Maybe GI, but when the author has trouble articulating basic things they lose credibility and I lose interest.