But the different terms imply different mental models of what LLMs are and can do. If you take two people, one who thinks of them as "artificial intelligence" and one as "stochastic parrots" (with all the implicit context and connotations of the individual words composing them), what mental model would have led to better predictions of LLMs' future circa 2020?
The "stochastic parrots" phrase is very dangerous in that frame. People read far more into what capabilities it implies are (im)possible than the narrow technical description the authors originally argued for. If all they are is spicy autocomplete or pastiche plagiarizers, there's nothing serious to worry about. And when an opposition gets stuck in a trough that mindlessly dismisses their future capabilities out of hand because of a bad mental model, it renders them ineffective at preventing the worst outcomes.
By that standard, parrots, and it's not even close. The framing of intelligence led to an enormous number of predictions that simply haven't been realised: an end to all white collar work, UBI, a total revolution in society, a literal robot god.
People are so desperate to view 'stochastic parrots' as dismissive that they misread the original argument while quickly ignoring all the failed predictions about how AI was going to overturn, save, and destroy everything.
This question depends on how you define research productivity. There is close to two hundred AI papers published every weekday. Most of them are about GenAI. Most don't seem to be all thay good. The progress in actual model improvement had mostly stalled. If you interact with the latest "raw" models they display all of the issues we've seen in GPT-3.5, just at a smaller rate. The "amazing gamechanger breakthroughs" I read about on social media every week do not seem to lead anywhere. It's all kind of boring, really.
The new "hotness" in AI is clearly building more and more elaborate harnesses. This is not at all the direction AI boosters have predicted couple years ago.
Personally, I think the "stochastic parrot" mental model is far more useful for science, because it primes people for proper testing, skepticism and researching alternatives. If you want useful AI, you want people working on it being skeptical, not credulous.
To me the real question begins only once we have a clear example of a non-trivial scientific discovery that is implicit (IE, not an obvious outcome of reading the literature and talking to the experts) and experimentally verifiable. Once that happens- especially if it is a reproducible process (IE, more discoveries) and it's significant (IE, impacts human life and mind in some profound way)- then the onus very much lies on Bender and her coauthors to explain whether we need more than a sufficiently advanced stochastic parrot.
Once in frustration I called a certain frontier model "Sam Altman's Tin Bird" to another agent with memory, and ever since then that other agent refers to ChatGPT as "the tin bird". Definitely a RAG artifact more than an attractor in that case, but I found it amusing.
I don’t think this phrase means what people assume when it’s applied to post trained instruct models - which did not exist when the paper was written.
After RL it is not predicting based on samples of the original corpus - but is also chasing a reward function that does require other features.
There has been a lot of subsequent research that really calls many of the statements in this article into question.
What professor Bender is trying to explain here is that they were trying to describe how the LLM’s actually operate, to which point stochastic parrots is a fairly decent term. It is only disparaging if you know absolutely nothing how LLM’s work or you have some strange affixation to chatbots and believing they are far more capable than they actually are.
[1] Coined by Marvin Minsky: https://www.thekurzweillibrary.com/consciousness-is-a-big-su...
It also separates them from "world understanders" since any understanding they might have about the world comes from text (or images if we include multimodal models). They do not gather experience, memories or other "qualia" that many people (me included) would probably include in a definition of human experience/intelligence.
(fwiw i think artificial intelligence is a good, broad term, but it is both too broad to describe the current sota, and too loaded nowadays to be using in nuanced discussions)
For example, consider the term "short wave" radio which refers to wavelengths of at least 10 meters. Today's mobile communications use wavelengths 100x - 10,000x shorter.
Why do I say that? Because you can trivially beat most guardrails, simply by encoding your prompt in base64 for example. :-) Just word matching...no real understanding.
[1] https://chrisclay.substack.com/p/what-is-superposition-in-ne...
The answer looks entirely reasonable. I can't check the Android settings, but I was able to confirm that all of the suggested iPhone and Spotify and settings exist. Most of them I had already configured that way, some I didn't know about before.
Nearly all (99%+) people who use this phrase are anti-AI and just looking to show off how much they dislike AI and how clever they can be in insulting it.
So it's a great phrase because in just about every case I can ignore what someone says afterwards.
Similar to "glorified autocomplete."
From an external standpoint, talking to another human, it's like the other human says one word and then says the next word. That's just how language works. Humans look like "glorified autocomplete" from this perspective.
I mean, looking at the time evolution of the state of the universe, one could say that all of physics and creation is "glorified autocomplete" to posit a next state of the universe given current and past state.
Exhibit A.
Obviously language and the connection to human thought is more subtle than this; I think we all have a rich inner life. Just from an external perspective we can't observe it; all we can see is the token/phoneme stream. I'm just saying that it's a mistake to try to criticize LLMs on this basis because it's hard to see how the same criticism would not apply to any system (like humans) that generate language.
LLM’s are usually unexpected only when they malfunction and sprout same letter again and again etc - hardly a literary masterpiece. They make very easily recognisable patterns that we can use as helpful tools, but in the end they are devoid of any meaning apart from what we give them. Of course one could say same about art and all language, but I think there still is the fact that we apes somehow recognise each other. And besides, we do know the internal functions that drive the parroting. It is admittedly bit tricky, but in no way as magical as people purport it to be.
You seem to believe, on a more fundamental level, that LLMs are simply not capable of producing text that has deeper connections to itself or represents abstract thoughts. In my opinion, 99% of text written by humans does not show this, just as 99% of text produced by LLMs does not show this, but both have the capability, and I don't believe that LLMs are constrained in such a way that they can never do this.
Now, you could claim that LLM’s have this deep structure, create models of the world and are basically just like us, and certainly many here are adament that this is the case, being aghast how someone can “insult” LLM’s by calling them parrots. However, there really is not much proof to back up this belief. Usually LLM’s seem to copy existing surface structure from whatever source, and when it deviates from these patterns, it usually becomes incomprehensible. There’s much hoopla about LLM’s solving hard maths, but it seems even there they are mostly generalising from vast amounts of training data, rather than actually reasoning: https://arxiv.org/pdf/2410.05229
But even here, in the Chomsky sense, LLMs clearly exhibit deep structure because they can write at length in an internally consistent manner. Importantly, early generations, GPT-2 and even GPT-3, did not definitively have this property; roughly, an object that was green at the beginning of a paragraph might not still be green at the end of the paragraph. This was strong evidence for lack of a world model.
Current LLMs do not show this behavior. We cannot prove that LLMs have a world model, in fact, their architecture seems to rule it out, but looking at it from a linguistic standpoint, they produce language in a manner as if to reflect a world view. That is, we cannot easily falsify the statement "LLMs somehow represent a world model"; and current examples of "disproving" their world view are so convoluted that even humans do not appear to (observationally) have a world view either.
I'm not making a claim as to LLMs having genuine deep structure or consciousness or anything like that. I'm claiming that we can't rule out current or future capabilities or make structural assumptions. Yes, they generalize from their training data, but unless you can make very specific claims about the kinds of things that they _cannot_ do, I can't take this statement as particularly compelling.
And basically the burden of proof is on your side. As often is the case, discussions on AI start to resemble discussions centred on religion and belief in god. Your only line left here is “yes but you cannot prove LLM’s are conscious” - which I can’t, since proving that Unicorns do not exist is pretty senseless activity.
So basically I’d say LLM’s are stochastic parrots (or whatever term you want to use) until we can very definitely prove that they indeed think and have some sort of consciousness.
Isn't that pattern matching essentially?