Modern language models refute Chomsky’s approach to language
scholar.google.com
scholar.google.com
The fact that you can replicate coherent text from probabilistic analysis and modeling of a very large corpus does not mean that humans acquire and generate language the same way. [edited page = 15]
Also, the LLMs are cheating! They learned from us. It's entirely possible that you do need syntax/semantics/sapience to create the original corpus, but not to duplicate it.
Let's see an AlphaZero-style version of an LLM, that learns language from scratch and creates a semantically meaningful corpus of work all on its own. It's entirely possible that Chomsky's mechanisms are necessary to do so.
No...they aren't. Humans aren't learning from thin air by any stretch of the imagination.
This is a complete tautology in terms, and any dispute must therefore be over the ontology in which it is expressed. To elaborate how something came to be is not required to express that something is.
Evolution optimized for language learning abilities before the full capabilities came. Evolution is not thin air.
I like to think of the brain as the general model we've trained with evolution, and the person's experiences as the specialization.
LLMs cheat at generating text because they do so via a model of the statistical structure of text.
We're in the world, it is us who stipulate the meaning of words and the structure of text. And we stipulate new meanings to novel parts of the world daily.
What else is an 'iPhone' etc. ? There's nothing in `i P h o n e` which is at all like an iphone.
We have just stipulated this connection. The machine replays these stipulations to us -- it does not make them, as we do.
It’s especially hard to parse a dark sweeping condemnation based on…people are investing in it? It doesn’t have the right to assign names to things? Idk what the argument is.
My most charitable interpretation is “it cant reason abour anything unless we already said it” which is obviously false.
The point is that they're not those things. Yes, language models can produce solutions to language tests that a 14 year old could also produce solutions for, but a calculator can do the same thing in the dimension of math - that doesn't make a calculator a 14 year old.
I removed the reference, in retrospect, it’s unnecessary. No need to indicate the strong performance, we’re all aware.
The point is not so much that we already said it, more that the patterns it encodes and surfaces when prompted are patterns in the written corpus, not of the underlying reality (which it has never experienced). Much like a list of all the addresses in the US (or wherever) will tell you very little about the actual geography of the place.
You've never experienced the "underlying reality" either.
No we don't. Humans don't experience or perceive reality. We perceive a nice modification of it and that's after excluding all sense data points we simply aren't capable of perceiving at all.
Your brain is constantly shifting and fabricating sense data based on internal predictions and that form the basis of what you call reality. You are not learning from the structure of the world. You are learning from a simplified model of it that is fabricated at parts.
Consider two people - one, a Papau New Guinea tribesperson from a previously uncontacted tribe who is allowed to handle a powered-down iPhone, and told it is an "iPhone", but is otherwise ignorant of its behavior - the other, a cross-platform mobile software developer who has never actually held a physical iPhone, but is intimately familiar with its build systems, API, cultural context etc. Between the two of them, who better understands what an iPhone "is"?
You make a good point about inventing words to refer to new concepts. There's nothing theoretically stopping a language model from identifying some concept in its training data that we don't have a word for, inventing a word for it, and using it to give us a perspective we hadn't considered. It would be very useful if it did that! I suspect we don't tend to see that simply because it's a very rare occurrence in the text it was trained on.
A concept is a sensory-motor technique abstracted into a pattern of thought developed by an animal, in a spatio-temporal environment, for a purpose.
LLMs are just literally an ensemble of statistical distributions over text symbols. In generating text, they're just sampling from a compressed bank of all text ever digitised.
We aren't sampling from such a bank, we develop wholey non-linguistic concepts which describe the world, and it is these which language piggy-backs on.
The structure of symbols in a book has nothing to do with the structure of the world -- it is we who have stipulated their meaning: there's no meaning to `i`
Text is an extremely limited input stream, but an input stream nonetheless. We know that animal intelligence works well enough with any of a range of sensory streams, and different levels of emphasis on those streams - humans are somehow functional despite a lack of ultrasonic perception and primitive sense of smell.
And your definition of a concept is quite self-serving... I say that as a mathematician familiar with many concepts which don't map at all to sensory motor experiences.
Sensory-motor expression of concepts is primitive, yes, they become abstracted --- and yes the semantics of those abstractions can be abstract. I'm not talking semantics, i'm talking genesis.
How does one generate representations whose semantics are the structure of the world? Not via text token frequency, this much is obvious.
I dont think the thinnest sense of "2 + 2 = 4" being true is what a mathematician understands -- they understand, rather, the object 2, the map `+` and so on. That is, the proposition. And when they imagine a sphere of radius 4 containing a square of length 2, etc. -- I think there's a 'sensuous, mechanical, depth' that enables and permeates their thinking.
The intellect is formal only in the sense that, absent content, it has form. That content however is grown by animals at play in their environment.
Hi, since human linguistics is the sole repository of linguistic conceptualism, can you please show me which of the neurons is the "doggie" neuron, or the "doggie" cluster of neurons? I want to know which part of the brain represents the thing that goes wag-wag.
If you can't mechanically identify the exact locality of the mechanism within the system, it doesn't really exist, right? It's just a stochastic, probabilistic model, humans don't understand the wag-wag concept, they just have some neurons that are weighted to fire when other neurons give them certain input stimuli tokens, right?
This is the fundamental problem: you are conflating the glue language with the implementation language in humans too. Human concepts are a glue-language thing, it's an emergent property of the C-language structure of the neurons. But there is no "doggie" neuron in a human just like there is no "doggie" neuron in a neural net. We are just stochastic machines too, if you look at the C-lang level and not the glue-language level.
But then also consider the following: a human being from 2006, and an LLM that has absorbed an enormous corpus of words about iPhones that is also granted access to a capacitive-touchscreen friendly robot arm and continuous feed digital camera (and since I'm feeling generous, also a lot of words about the history and architecture of robot arms and computer vision). There is no doubt the LLM will completely blow the human out of the water if asked trivia questions about the iPhone and its ecosystem.
But my money's on the 2006 human doing a lot better at switching it on and using the Tinder app...
This just comes down to giving transformers more modalities, not just text tokens.
There is nothing about “2” that conveys any “twoness”, this is true of all symbols.
The token “the text ‘iphone’” and the token “visual/tactile/etc data of iphone observation” are highly correlated. That is what you learn. I don’t know if you call that stipulation, maybe, but an LLM correlates too in its training phase. I don’t see the fundamental difference, only a lot of optimizing and architectural improvements to be made.
Edit: and when I say “a lot”, I mean astronomical amounts of it. Human minds are pretty well tuned to this job, it’ll take some effort to come close.
> Humans learn from the structure of the world -- not the structure of language.
You'd be surprised. Many researchers believe that "knowledge" is inseparable from language, and that language is not associative (labels for the world) but relational. For example, in Relational Frame Theory, human cognition is dependent on bidirectional "frames" that link concepts, and those frames are linguistic in nature. LLMs develop internal representations of those frames and relations, which is why they can tell you that a pool is bigger than a cup of water, and which one you would want to drink.In short, there's no evidence that being in the world makes our knowledge any different from an LLM. The main advantages we have at the moment are sensory learnings (LLMs are not good at comparing smells and flavors) and the ability to continuously train our brains.
It almost doesnt matter what your theory of language is --- any even plausible account will radically depart from the above statistical model. There isn't any theory of language which supposes it's an induction across text tokens.
The problem in this whole discussion is that we know what these statistical models are (models of association in text tokens) -- yet people completely ignore this in favour of saying "it works!".
Well "it works" is NOT an explanatory condition, indeed, it's a terrible one. If you took photographs of the night sky for long enough, you'd predict where all the stars are --- these photos do not employ a theory of gravity to achive these.
LLMs are just photographs of books.
There's a really egregious pseudoscience here that the hype-cycle completely suppresses: we know the statistical form of all ML models. We know that via this mechanism arbitrarily accurate predictions, given arbitrarily relevant data, can be made. We know that nothing in this mechanism is explanatory.
This is trivial. If you video tape everything and play it back you'll predict everything. Photographing things does not impart those photographs the properties of those things -- those serve as a limited assocative model.
I don’t have a reference handy now (someone can probably do better) but I believe one way to see this is via the hearing impaired or hearing and sight impaired
Fine.
We know that languages did not exist at some point. We made them exist.
Whether that happens in 1 generation or 10,000 generations is completely unknown.
> Steven Pinker, author of The Language Instinct, claims that "The Nicaraguan case is absolutely unique in history ... We've been able to see how it is that children—not adults—generate language, and we have been able to record it happening in great scientific detail. And it's the only time that we've actually seen a language being created out of thin air."
Humanoids went millions of years literally learning to navigate 3D space and sense “enough heat, food, water” etc
Nomadic tribes had built shared resource depots millennia before language.
I can see the color gradients of the trees and feel muscles relax without words.
Human language beyond some utilitarian labels just instills mind viruses that bloom into delusions of grandeur.
90% of human communication is unspoken. Neuroscience shows our brains sync behavior patterns with touch and just being in a room.
Reality is full of unseen state change every moment that we have no colloquial language for; human language is hardly the source of truth and the “North star” of human society in reality.
This is as scientific as the idea of humans just using 10% of our brains.
There’s just as little science language motivates me to work. Most of the language society relies on is hallucinations; fiat currency, nation states, constructs like “Senate” and Congress, corporatism, brands, copy-paste of historical terminology, not evidence they’re immutable features of reality.
What we recite has nothing to do with what we are. I find the appeals to non-existent political truisms primate gibberish.
It seems perfectly clear to me many facts of society are just memorized and recited prompt hacks. Language is the goto tool for propagandists, to obscure sensory connection to reality.
There is over 100 years of propaganda research available, too much for me to sort through, but scientific measure of such is not new; new to anyone unaware of it but not to humanity.
Just because it intuitively makes sense doesn’t make it correct or the 90% figure accurate.
We can also look at creole languages. When distinct linguistic groups come into contact without a common language they will, in some cases, develop "pidgin" languages to communicate. These languages are created by the adults who are already fluent speakers of their native language. What is interesting is what happens with children born into such communities. They grow up hearing the pidgin language, but they do not learn the pidgin. Instead, the language they learn is a creole. It is based on the pidgin, but more complete and consistent with the way human language works in general. In effect, like all children, they are relying on their innate knowledge of language. The pidgin fills in the language specific parameters, but when the pidgin contradicts an aspect of innate language the children (for the most part), just don't aquire that portion of the language.
LLMs are not doing this.
LLMs need pairwise linear input that is composed of independent and identically distributed data.
Feed forward neural networks are effectively DAGs thus semi-decidable.
LLM require a corpus, that data is generated by humans and isn't a perfect information game.
If you dig into how many feedforward neural network can be written as a single pairwise linear function in lower dimensions, you can help build an intuition on how they work in higher dimensions that are beyond our ability to visualize.
AlphaZero being able to build a model without access to opening books or endgame tables in perfect information games was an achievement in implementation, it was not a move past existential quantifiers to a universal quantification.
LLMs still need human produced corpus because the search space is much larger than a simple perfect information game. The game board rules were the source of compression for AlphaZero, while human produced text is the source for LLMs.
Neither have a 'common sense' understanding of the underlying data, their results simply fit a finite subset of the data in the same way that parametric regression does.
As there are no accepted definitions for intelligence, mathematics is the only way to understand this.
VC dimensionally and set shattering is probably the most accessible to programming backgrounds if you are interested.
Ding ding ding, right answer!
Now I leave you with a thought: given that the AI is learning from us, is it possible to have an AI that's much more intelligent than their "source of truth"?
Sure, I can invent a new language for you! Let's call it "Vorin" for the purposes of this demonstration.
Vorin is a tonal language with a complex system of noun classes and a relatively simple verb conjugation system. It is spoken by the Vorin people, who live in a mountainous region of an imaginary world. Here is a sample conversation in Vorin:
Person 1: Ní lé màrì?
Person 2: Lé vàlà, nì síkòní.
Person 1: Vàlà síkòní?
Person 2: Àmà, síkòní là nì tékàlá.
Translation:
Person 1: How are you?
Person 2: I'm fine, thank you.
Person 1: Are you working?
Person 2: Yes, I'm working right now.
In Vorin, the tone of each word is very important, and can change the meaning of the word entirely. For example, "lé" with a high tone means "you", while "lé" with a low tone means "fine". In addition to tone, Vorin has a system of noun classes, which are indicated by prefixes on the nouns themselves. For example, "màrì" means "you" if it is in the second person singular class, but it can also mean "house" if it is in the third person singular class.
Vorin's verb conjugation system is relatively simple, with only a few tenses and moods. Verbs are also marked for subject agreement, which means that the verb ending changes depending on the subject of the sentence. For example, "nì síkòní" means "I'm fine", while "là síkòní" means "he/she/it is fine".
Overall, Vorin is a complex and interesting language with a rich grammatical structure and unique tonal system.
Let's see it invent language from scratch without being presented any examples of languages to use as training data. That's what humanity did. Can LLMs do it? I doubt it, not yet.
1."lè" with a low tone means fine, but "nì síkòní" means "I'm fine", while "là síkòní" means "he/she/it is fine"?
2. The example of "noun prefixes" doesn't include different noun prefixes?
3. The example of "verb ending changes" doesn't include different verb endings?
4. Since the second clause of the second phrase apparently means "I'm fine", the first word must mean "you", but then the second word is repeated in the next sentence which has nothing to do with thanking someone? And "síkòní" gets sprinkled into three sentences with no common concepts?
This just looks, predictably, like gibberish spit out by an LLM.
Maybe the LLM should write the tones too, so it sounds like a child inventing a language and when you point out its logical inconsistencies it invents new rules to fit.
The results were... not all that impressive. There were significant issues getting it to consistently apply the rules of the language it had created, even from one prompt to the next -- and after a certain point, it decided to just give Arabic translations instead of the conlang it was supposed to be making up.
Perhaps a more dedicated "prompt engineer"/linguist-type might be able to get better results, but the problem here seems to be similar to the problem trying to get ChatGPT to do arithmetic and other extreme sports. When trying to get it to do anything other than generating one-off syntactically-correct responses to simple prompts in already-existing human languages, it falls down horribly.
They do seem to need a significantly larger corpus, though, so it's not clear that it actually refutes Chomsky.
I am not sure how brute forcing a chess using Monte Carlo Tree Search, or solving Checkers via exhaustive search, would refute a theory about how people with efficient, low-power-consumption brains that grow organically, are able to master Chess.
Obviously, all that stuff ChatGPT says about feelings and emotions came from humans writing it!
Therefore, if we're able to do the same thing by simply applying more resources, that would undermine his argument in a way that doing the same thing with a vastly larger corpus (whatever the resources we throw at it) doesn't.
I should note that this is based on recollections from 20+ years ago and no serious engagement with the article at hand, so, uh, appropriate salt.
The simple reason is: because they don't actually "know better". Maybe they are knowledgeable and skilled in some area, but that doesn't mean they are knowledgeable and skilled in everything.
(I’m agreeing with you basically)
As for something LLMs are unlikely to do under any circumstances, there's already a fairly obvious example. They can't keep a secret, hence prompt injections.
I'm less convinced there's any unified solution for "general purpose AI" before us here.
> Wordcels are people who have high verbal intelligence and are good with words, but feel inadequately compensated for their skill. The term "cel" denotes frustration over being denied something they feel they deserve.1 Shape rotators are people with high visuospatial intelligence but low verbal intelligence, who have an intuition for technical problem-solving but are unable to account for themselves or apprehend historical context.2 The use of the terms has skyrocketed online in the past few months, especially in the last few days.0 The term "wordcel" is derived from incel and is used to describe someone who has high verbal intelligence but low "visuospatial" intelligence, whose facility for and love of complex abstraction leads them into rhetorical and political dead-ends.
Biology is intrinsically local. For Chomsky’s model of language instinct to work, it would have to reduce down to some sort of embryonic developmental process consisting of entirely of local gene-activated steps over the years it takes for a human child to begin speaking grammatical sentences. This is in direct contrast to most examples of human instinct, which disappear very quickly as the brain develops.
Really the main advantage that Chomsky’s ideas had is that no one could imagine how something simpler could possibly result in linguistic understanding. But large language models demonstrate that no, actually one simple learning algorithm is perfectly sufficient. So why evoke something more complex?
Zero to one is closer to mimicry and immersion. There's a long Wikipedia article on the field of study https://en.m.wikipedia.org/wiki/Language_acquisition
Furthermore, humans probably aren't static learners and likely have more beneficial times of certain study than others. There's a theory in that too https://en.m.wikipedia.org/wiki/Critical_period_hypothesis
Saying there's a "digital brain" is more of a framework since the term "brain" looks like it's a moving target
In another comment I referred to these systems as like comparing hydraulic pumps to human biceps, cars to horses, etc.
We can use the same units of measure, give them the same tasks, but saying they're the same thing only works in the world of poetry
If you give it a small enough training set or a big enough neural network, it will directly memorize the whole thing. You have to intentionally make its brain too small to do that in order to force it to find patterns in the data instead.
These types of "mistakes" are more about the authors letting their intentions and hopes known on how they wish the thing to be used.
Yeah the whole thing hinges on this... and uh yeah good luck with that one...
The Norvig-Chomsky debate is kind of old at this point:
https://www.tor.com/2011/06/21/norvig-vs-chomsky-and-the-fig...
Chomsky is trying to explain how humans create language. LLM are creating language, but not the way humans do.
Nothing about this paper refutes Chomsky's claims.
In neuroscience, predictive processing has gained immense favor and can explain language in ways that have nothing to do with innate grammar.
https://en.wikipedia.org/wiki/Predictive_coding
Exactly how well did "building a bird" work for building flying machines? Birds use the same principle as a fixed wing when it comes to soaring flight. "Building a bird" without the principles of an airfoil and just mimicking the flapping wings does not result in flight.
Re: predictive processing, in what way does that relate to language…? Even if you apply it to language in a way not mentioned in the linked article at all, I don’t see even the rough shape of how it would refute (/be mutually exclusive with) generative grammars. Maybe I’m just missing something because I don’t know much neuro?
Re: building a bird… yeah that’s their point, you don't try to build a bird, you try to study birds. Chomsky cares about what we are, not building machines to do our drudgery. I don’t think I agree entirely with that singular focus, but you see the appeal, no?
https://projects.iq.harvard.edu/kuperberglab/people/gina-r-k...
Here's some relevant research:
https://projects.iq.harvard.edu/kuperberglab/publications/dy...
I'll point out that neuroscientists have yet to find the "generative grammar" part of the brain but have seen evidence of a very large network of neurons...
But I can’t help by think about all of the whacky ideas like antigravity vital forces that biologists contrived to explain how birds could fly and that it took Bernoulli and the rigorous study of those principles that led to the airfoil… which is how birds actually soar through the air.
BTW, what the fuck happened to these forums in the last few years? It seems like most people base their opinions on how opposite they are to Sam Altman and Elon Musk’s as opposed to any geeky principles of discovery. I highly doubt that most of y’all would have been ardent supporters of generative grammar five years ago… but slap the word LLM on something and boy howdy!
It’s kind of nice, I learn quite a bit defending good ideas. All y’all get is fake internet points.
I fully understand any and all downvotes! Have fun!
Biologists don't look at a 737 and say "that's obviously the way flight works, I wonder where the engine on that seagull is".
https://en.wikipedia.org/wiki/Generative_grammar
https://en.wikipedia.org/wiki/Transformational_grammar
https://en.wikipedia.org/wiki/X-bar_theory
https://en.wikipedia.org/wiki/Principles_and_parameters
Not really. You could probably put together a flapping bird toy (can buy these mass manufactured too) in about a couple of months of trial and error. Not quite as sophisticated as feathers but the principles are the same. You probably couldn't build an airplane.
This paper: Planes fly, but don’t flap their wings, ergo Chomsky is wrong.
we actually don't know what is inside LM too, so it is possible LM statistically learns syntax and semantics, and it is major part of output quality.
LLMs can code in the same way they can use natural languages. But we know that programming languages have structure, we made them that way, from scratch, using Chomsky's theory no less.
Saying that because LLMs can learn programming languages using a different approach and therefore disprove the very theory they are built on is absurd.
Anyways, the paper is long and full of references, I didn't analyse it, does it include looks inside the model? For example, for LLMs to write code correctly, the structure of programming languages must be encoded somewhere in the weights of the model. A way to more convincingly disprove Chomsky's ideas would be to find which part of the network encodes structure in programming languages, and show that there is nothing similar for natural languages.
Very much so, it's astounding really. I still remember deriving "words" and using Chomsky's Normal Form when making the CFG to build a compiler.
We have known that it is possible to understand language without innateness. That is what linguists do.
If you look at how linguists know about innate features, the answer is almost always by first discovering them explicitly while analysing language data; not by opening a brain to see what is innately inside. [0]
The point about innateness is that it takes generations of linguists to learn from a blank slate properties of language that children learn in just years.
There are also numerous other arguments for innateness. From the way humans seem to spontaneously develop language in a language deprives environment, to the way language aquasition works being more consistent with other innate behaviours, to the pressence of weird properties that seem to be present across languages for no apparent logical reason.
The only insight I see from LLM is the same insight we have seen throughout macine learning. It is not nessasary to understand something if you can throw enough compute at it. This is powerful, and it enables us to do a lot, but it should not be confused with understanding.
[0] There are some instances leveraging MRI and other cognative research teqniques to get some insight into the inner workings of human language processing, but their role in developing current linguistics theory is thus far limited.
"Language is innate in humans because in every household practically all children learn it while none of the pets do."
[1]: at least the mamal pets, not goldfishes
We teach children to read and write, but we don't need to teach them to listen and speak.
And whales might say something similar about us, that humans have a very complex society despite being seemingly unable to learn to speak coherently.
On a lighter note, I do expect “Modern Language Models Refute…” to be the new “All you need is…”! It’s just too provocative not to click on
This out-of-the-blue accusation sounds like a confession of your true motives in this conversation: You like the man's politics, so you feel compelled to defend him in an unrelated topic.
> First, the fact that language models can be trained on large amounts of text data and can generate human-like language without any explicit instruction on gram- mar or syntax suggests that language may not be as biologically determined as Chomsky has claimed. Instead, it suggests that language may be learned and developed through exposure to language and interactions with others.
I'm not a linguist nor a cognitive scientist, but this seems so problematic that I am not sure that I read it correctly. For example, how is the fact that language models "work" contradict the innateness of language in humans?
(Similarly, languages that have gender are typically just picked up by usage, not necessarily ingrained by reasoning. Which leads to the obvious bad results when people think that there was solid reasoning on those choices, in the first place.)
how much is innate, what exactly does that mean, all good questions. and of course raw intelligence (pattern matching, strategizing, learning, adaptiveness, modeling, ability to form a sort of consistent and goal-orientedly useful predictive model of the world based on inputs, and goal-oriented control of behavior based on these aforementioned models) by definition can learn language.
and of course it's a strange question of are LLMs intelligent in this sense despite lacking goals?
It doesn't. Also, the author doesn't seem to actually understand Chomsky's writing about language, because learning language via exposure is how humans learn languages and he fucking mentions that in his writing on the subject.
UG (universal grammar) is the purported facility in the human brain which makes language possible - it has a innate structure, but it learns particular languages from exposure. Chomsky doesn't state exactly what that structure is because he doesn't know - figuring that out is goal of his work.
How much input data is used to train modern language models?
The combinatorial space of languages is obviously infinite, but is this so for grammars? If not then you would expect many languages to share the same grammars.
Seems analogous to the same question about mathematics. Are there many different possible arithmetics? No. There are infinitely many ways to express arithmetic symbolically but there is only one arithmetic. 2 + 2 never equals 5.
LLMs started getting interesting "all the sudden" when they hit a certain scale, just like biological brains.
There's lots of animals that have very complex brains, and that engage in very complex behaviors - the fact that none of them have even a hint of linguistic ability seems like a strong indicator that these faculties are a human-specific evolution, and not just... well IDK how to even sum up an anti-UG/GG view. "Kids just sorta figure it out" I guess?
Although who knows! As Chomsky likes to say, this whole field is in a pre-Galilean state due to the impossibility of conducting comparative studies
EDIT: Oh just realized you were the parent comment. Well I'd say the small NN vs. LLM example still doesn't convince me of the likelihood of your statement as I understand it; it goes without saying that lots of animals are much better at intuitive understandings of physics (well, kinetics at least) than humans. You ever seen those snakes that jump from tree to tree? craziest shit you'll ever see
2. Depending on what exactly you're trying to teach (perfect grammar, paragraphs of coherent text, basic reasoning), much less data is needed. https://arxiv.org/abs/2305.07759
3. Brains don't start at 0. Evolution, dna/rna etc. There's obviously some pre disposition for language learning in humans but that alone isn't enough ground for a "universal grammar"
4. We really do take in an enormous amount of data (not text specifically)
Humans have some 50 to 100 trillion synapses. Who knows how far scaling goes in a transformer but simply increasing the parameter count increases performance so far.
Either way, we're not really close to emulating the complexity of the brain especially when taking neurons into account (one human neuron is far more complex than any one artificial parameter)
In other words, simpler building blocks and smaller sie. Of course, how much complexity is needed for Intelligence is unknown.
That’s at least a quadrillion parameters which is at least three orders of magnitude bigger than SOTA LLMs assuming each of those synapse channels maps to a parameter. That is an absurd assumption given neuroplasticity, which operates at a high level to adapt the neural network as it learns. See the plasticity section in the linked wikipedia articles: connections between neurons can grow or even get removed on the scale of hours and days. The human brain has entire biological systems supporting intelligence for which there is no real ML equivalent because our ML architectures are static.
Given a family member’s research on neurotransmitter potentiation, I’d estimate 10 to 100 parameters per channel to get full fidelity of the human brain. (This is absurdly speculative - we have no idea at which fidelity will intelligence emerge)
If GPT-4 has 100 trillion parameters, it has as many parameters as the human brain has synapses. Synapses are a lot simpler than parameters; they're digital. A single neuron needs many synapses, all of roughly equal weight, emitting many pulses over a short time in order to convey a single weighted value.
On top of that, you may have heard that the human brain does a lot of things besides writing. You subtract the motor cortex, the visual and limbic systems etc... a 100 trillion parameter model is unambiguously larger than the language processing portions of the human brain.
> Brains don't start at 0. Evolution, dna/rna etc. There's obviously some pre disposition for language learning in humans but that alone isn't enough ground for a "universal grammar"
The human genome is 24 gigabits long. It's negligibly small compared to a language model.
It doesn't.
>Synapses are a lot simpler than parameters
Not true but someone else has explained why.
>a 100 trillion parameter model is unambiguously larger than the language processing portions of the human brain.
GPT-4 is not that big. The technology to run such a model at scale is simply not feasible yet.
>The human genome is 24 gigabits long. It's negligibly small compared to a language model.
I'm sorry bit this makes no sense. How many gigabits long the genome is has no bearing on how impactful it is in steering the development of the human brain in comparison to a ML model.
That is literally what universal grammar is. All that is left is to argue about the size and content of UG.
https://www.scientificamerican.com/article/evidence-rebuts-c...
20 year old human has
* heard ~220 million words, talked 50 million words.
* read ~10 million words.
* experienced 420 million seconds of wakeful interaction with the environment (can be used to estimate the limit to conscious decisions, or number of distinct 'epochs' we experience)
From a machine learning perspective human life is surprisingly small set of inputs and actions, just a blip of existence.
But also humans have been speaking for so long it's silly to imagine we don't have some evolved language structures in the brain. I don't know why anyone would single that out for skepticism while not questioning e.g. the brain structures for sight, sound, emotions, navigation, etc.
That's Chomsky's argument. A small set of constraints for organizing language.
Put a kid from one language tradition in a spot with a different language tradition, and they won't be able to learn it.
Eg. kids with native mandarin speaking parents adopted to native Indo-European parents fail at learning English, and will be better at learning Mandarin than their peers with Indo-European heritage.
[1] https://www.scientificamerican.com/article/evidence-rebuts-c...
That work fails to support Chomsky’s assertions. The research suggests a radically different view, in which learning of a child’s first language does not rely on an innate grammar module.
Instead the new research shows that young children use various types of thinking that may not be specific to language at all—such as the ability to classify the world into categories (people or objects, for instance) and to understand the relations among things.
These capabilities, coupled with a unique human ability to grasp what others intend to communicate, allow language to happen.
The fact that very smart people think this refutes Chomsky makes me quite sad. They basically restated the UG theory in the last sentence, as proof that it’s wrong…Chomsky has been saying for literal decades that language is likely a corollary to the basic reasoning skills that set humans apart, but people still think UG means “kids are born knowing what a noun is” :(
This also reminds me of evolution. Some people looked at discoveries in epigenetics and declared that it disproved Darwinian evolution in favor of Lamarckian evolution.
Sure, Darwin's theory of natural selection combined with random variation at the point of reproduction does not explain 100% of evolution, but it is still covers most of it.
This is opposed to modeling a concept in your mind, and then applying language through denotation. This isn't unlike composing a request to be sent over a specific protocol. The data exists independently of the protocol and could even be fit to more protocols, with the right understanding of how to implement them. Sure, you have to read some docs, and maybe use a library somebody else wrote, but nobody in their right mind would call that plagiarism. This is more akin to how language works in the human brain, where each new language is like a different protocol.
"The fact that this advanced drone released in the year 2080 that can drive with the same agility as an eagle proves that the eagle's flying ability is not as biologically determined as some people claim.
In fact, any organism can fly if it sees enough data about flying!"
It is entirely likely that the way we operate is probability-first, only deriving rules loosely after taking in lots of experiential data to speed up and simplify that initial fake-it-til-you-make-it understanding. The fact LLMs can get the quality we see using just this approach is a strong indicator that this method of understanding may be a fundamental approach of biological systems too.
(and if you're arguing this is unfair because humans created the language that's being used for probabilistic training - well, look at image models trained on photographs instead and tell me those aren't an example of extreme quality derived purely from mass-inferenced data. Rules-based architectures don't necessarily need apply.)
But honestly, this seems like a silly claim to begin with if it really was claimed. We have formal language theory complexity classes of probabilistic algorithms for a reason - they work! It shouldn't be surprising that the model can stretch down to the fundamentals too. Far fewer programmers (and linguists) were raised to think with these models than deterministic rules-based ones, but the field has been progressing alongside for decades, and now they get to play with powerful LLMs that take probabilistic inferencing to the extreme and will likely prove it works (very elegantly) for everything. This shouldn't be surprising in retrospect.
Chomsky may very well be right that there always exists some fundamental elegant formula underlying any phenomenon (or at least any language). But it's undeniable at this point that simplistic statistical approaches can be applied at scale to those phenomenon and derive highly-useful general models, which also will very likely converge upon the elegant formulas he envisioned. The two are intrinsically linked, neither inseparable.
The author seems to be stating that because llms have generated "interesting" statements that all of Chomsky's approaches are invalidated. This seems to assume that the llm isn't somehow incorporating Chomsky's generative grammar system into the model, which would be a very hard thing to prove.
Also, the author's stated proof contains flaws. Take for example page 16:
Generate ten other sentences like "colorless green ideas sleep furiously": Purple fluffy clouds dream wildly. Blue glittery unicorns jump excitedly.
The author claims that the llm has completed this task successfully, completely ignoring the example's pairwise incompatible terms (colorless green, and sleep furiously), and accepting that "purple fluffy" is equivalently meaningless. It is not, and the model has clearly failed.
A computer has no sense of "intent" or "meaning" to what they have indexed and scanned. There is no intent or meaning to what they spit out in generated text.
We're back in the Chinese room. https://en.wikipedia.org/wiki/Chinese_room
Let's make sure what the computer scientists understand Chomsky to have stated is actually aligned. Chomsky didn't say the ONLY way to create language is via the brain. His view, instead, is that evolution programmed language development into the brain -- that it is not learned (entirely) by peer osmosis. That the brain has some structure for language, built-in, which is the unlocked in various ways via socialization.
Summary of Chomsky's view, paraphrased:
"It is a strange intuition [that most other people have]. Above the neck, we insist everything [in human development] comes from experience. Below the neck, we're willing to accept the idea that [...] it comes from inside. [...] But: it is hard to look at the Sun setting and say, it's not 'setting', the Earth is actually turning. Similarly, with people, it's hard for us to look at them and not see them as minds inside bodies. This leads us to this false approach: below the neck, we are willing to pursue the sciences, and if that leads us to believe development is internally programmed, we'll accept it. But above the neck, we'll be completely irrational; we're going to insist on beliefs and explanations we'd never normally dream of in other rational areas."
---
This YouTube clip came to mind, but here is a more detailed explainer from Stanford Encyclopedia of Philosophy:
"Clearly, there is something very special about the brains of human beings that enables them to master a natural language — a feat usually more or less completed by age 8 or so. ... This article introduces the idea, most closely associated with the work of the MIT linguist Noam Chomsky, that what is special about human brains is that they contain a specialized ‘language organ,’ an innate mental ‘module’ or ‘faculty,’ that is dedicated to the task of mastering a language.
On Chomsky's view, the language faculty contains innate knowledge of various linguistic rules, constraints and principles; this innate knowledge constitutes the ‘initial state’ of the language faculty. In interaction with one's experiences of language during childhood — that is, with one's exposure to what Chomsky calls the ‘primary linguistic data’ or ‘pld’ — it gives rise to a new body of linguistic knowledge, namely, knowledge of a specific language (like Chinese or English). This ‘attained’ or ‘final’ state of the language faculty constitutes one's ‘linguistic competence’ and includes knowledge of the grammar of one's language. This knowledge, according to Chomsky, is essential to our ability to speak and understand a language (although, of course, it is not sufficient for this ability: much additional knowledge is brought to bear in ‘linguistic performance,’ that is, actual language use)."
source: https://plato.stanford.edu/entries/innateness-language/
But does it really understand language? This reminds me of the Chinese room argument [1].
There is some research into this https://phys.org/news/2022-05-renormalization-group-methods-...
There is a fairly interesting read from Steven Pinker: "The language instinct" about this topic.
One of Chomsky's main arguments is the poverty of stimulus (children seem to learn language with relatively little input). Here is what the author has to say:
> Large language models essentially lay this issue [poverty of stimulus] to rest because they come with none of the constraints. Modern language models refute Chomsky’s approach to language that others have insisted are necessary, yet they capture almost all key phenomena. It will be important to see, however, how well they can do on human-sized datasets, but their ability to generalize to sentences out-side of their training set is auspicious for empiricism.
That doesn't look like a refutation to me yet. We still need to do that test, but that still just tells us how you can do it algorithmically.
Of all possible neural network architectures, so far one of them, Transformers, delivered good results.
It's possible this architecture is more similar to the brain language architecture.
"Only human can learn and understand language, not other creature in the university, not machine"