How much input data is used to train modern language models?
How much input data is used to train modern language models?
2. Depending on what exactly you're trying to teach (perfect grammar, paragraphs of coherent text, basic reasoning), much less data is needed. https://arxiv.org/abs/2305.07759
3. Brains don't start at 0. Evolution, dna/rna etc. There's obviously some pre disposition for language learning in humans but that alone isn't enough ground for a "universal grammar"
4. We really do take in an enormous amount of data (not text specifically)
Humans have some 50 to 100 trillion synapses. Who knows how far scaling goes in a transformer but simply increasing the parameter count increases performance so far.
Either way, we're not really close to emulating the complexity of the brain especially when taking neurons into account (one human neuron is far more complex than any one artificial parameter)
In other words, simpler building blocks and smaller sie. Of course, how much complexity is needed for Intelligence is unknown.
That’s at least a quadrillion parameters which is at least three orders of magnitude bigger than SOTA LLMs assuming each of those synapse channels maps to a parameter. That is an absurd assumption given neuroplasticity, which operates at a high level to adapt the neural network as it learns. See the plasticity section in the linked wikipedia articles: connections between neurons can grow or even get removed on the scale of hours and days. The human brain has entire biological systems supporting intelligence for which there is no real ML equivalent because our ML architectures are static.
Given a family member’s research on neurotransmitter potentiation, I’d estimate 10 to 100 parameters per channel to get full fidelity of the human brain. (This is absurdly speculative - we have no idea at which fidelity will intelligence emerge)
If GPT-4 has 100 trillion parameters, it has as many parameters as the human brain has synapses. Synapses are a lot simpler than parameters; they're digital. A single neuron needs many synapses, all of roughly equal weight, emitting many pulses over a short time in order to convey a single weighted value.
On top of that, you may have heard that the human brain does a lot of things besides writing. You subtract the motor cortex, the visual and limbic systems etc... a 100 trillion parameter model is unambiguously larger than the language processing portions of the human brain.
> Brains don't start at 0. Evolution, dna/rna etc. There's obviously some pre disposition for language learning in humans but that alone isn't enough ground for a "universal grammar"
The human genome is 24 gigabits long. It's negligibly small compared to a language model.
It doesn't.
>Synapses are a lot simpler than parameters
Not true but someone else has explained why.
>a 100 trillion parameter model is unambiguously larger than the language processing portions of the human brain.
GPT-4 is not that big. The technology to run such a model at scale is simply not feasible yet.
>The human genome is 24 gigabits long. It's negligibly small compared to a language model.
I'm sorry bit this makes no sense. How many gigabits long the genome is has no bearing on how impactful it is in steering the development of the human brain in comparison to a ML model.
That is literally what universal grammar is. All that is left is to argue about the size and content of UG.
https://www.scientificamerican.com/article/evidence-rebuts-c...
20 year old human has
* heard ~220 million words, talked 50 million words.
* read ~10 million words.
* experienced 420 million seconds of wakeful interaction with the environment (can be used to estimate the limit to conscious decisions, or number of distinct 'epochs' we experience)
From a machine learning perspective human life is surprisingly small set of inputs and actions, just a blip of existence.
But also humans have been speaking for so long it's silly to imagine we don't have some evolved language structures in the brain. I don't know why anyone would single that out for skepticism while not questioning e.g. the brain structures for sight, sound, emotions, navigation, etc.
That's Chomsky's argument. A small set of constraints for organizing language.
Put a kid from one language tradition in a spot with a different language tradition, and they won't be able to learn it.
Eg. kids with native mandarin speaking parents adopted to native Indo-European parents fail at learning English, and will be better at learning Mandarin than their peers with Indo-European heritage.
The combinatorial space of languages is obviously infinite, but is this so for grammars? If not then you would expect many languages to share the same grammars.
Seems analogous to the same question about mathematics. Are there many different possible arithmetics? No. There are infinitely many ways to express arithmetic symbolically but there is only one arithmetic. 2 + 2 never equals 5.
LLMs started getting interesting "all the sudden" when they hit a certain scale, just like biological brains.
There's lots of animals that have very complex brains, and that engage in very complex behaviors - the fact that none of them have even a hint of linguistic ability seems like a strong indicator that these faculties are a human-specific evolution, and not just... well IDK how to even sum up an anti-UG/GG view. "Kids just sorta figure it out" I guess?
Although who knows! As Chomsky likes to say, this whole field is in a pre-Galilean state due to the impossibility of conducting comparative studies
EDIT: Oh just realized you were the parent comment. Well I'd say the small NN vs. LLM example still doesn't convince me of the likelihood of your statement as I understand it; it goes without saying that lots of animals are much better at intuitive understandings of physics (well, kinetics at least) than humans. You ever seen those snakes that jump from tree to tree? craziest shit you'll ever see