Of course they have no understanding. They’re text generators.
This persistent response to LLMs is bigger news, to me, than the fact that LLMs can’t think.
Of course they have no understanding. They’re text generators.
This persistent response to LLMs is bigger news, to me, than the fact that LLMs can’t think.
I can testify that people will be manipulated by sleight of hand (the physical act) even after they have been alerted to it. Now transfer that weakness to a domain where most people have zero equipment and training to form a sound judgement.
Its a big and complicated world and our brains are wonderful but limited. We have to rely on trust for the vast majority of information we digest.
When organized coalitions of individuals subvert the channels of trust we stand exposed as idiots.
The principal characteristic of the digital era is, alas, not harnessing this amazing technology for societal benefit but betraying trust and violating unspoken contracts - for profit.
This can do that.
"The mode of action of paracetamol has been uncertain, but it is now generally accepted that it inhibits COX-1 and COX-2 through metabolism by the peroxidase function of these isoenzymes."
Furthermore, the active ingredient in Tylenol is a very simple molecule from a synthetic perspective, and it was developed while trying to overcome the toxicity of an even simpler molecule called acetanilide. Even though it was developed 150 years ago, there was a systematic understanding of organic chemistry and the interaction with the human body already at that time.
[1] https://link.springer.com/article/10.1007/s10787-013-0172-x
No? Maybe the analogy you are suggesting isn’t the most accurate.
These threads are always massive, embarrassing trainwrecks of naive takes on cognition and consciousness.
It's like HN just woke up and refuses to acknowledge the relevance of the extensive prior literature on these subjects. Instead, people seem mostly fine with making shit up.
Embarrassing.
It's fine. We're social animals, we discuss topics even if we lack perfect expertise, at our present level of understanding. I've certainly had, for example, many opinions on programming that were later superseded by deeper understanding or evolving priorities.
It's just important to remember there's probably a bigger and better expert out there, and stay humble and interested, and let yourself be corrected.
Among the other technical and scientific topics often discussed here, there is a basic general expectation of grounded reference to existing research and bodies of knowledge. Deviations from this standard are usually met with corrections.
With the topic of cognition and consciousness, this is not so (generally). The standard is to absolutely not reference existing research and to basically pretend it doesn't exist.
Instead, the vacuum of ignorance is filled with whatever comes to mind and threads engage in freeform speculation.
This is also ironic since these topics are usually about how LLMs don't understand and just fill the vacuum of its ignorance with whatever happens to be next on the statistical chain, regardless of veracity.
With physics, for example, people have an understanding that there's a large existing body of work, and that physics has been remarkably successful at building on earlier knowledge for decades without having to debunk a lot of earlier results. It's a familiar story.
With AI, however, the popular narrative is that earlier approaches to the problem were all laughing stocks that went nowhere. There's not a good sense for how old some the ideas that work now actually are, and how many of the older theoretical and thought frameworks are still perfectly valid, in particular at the interfaces toward other branches of science and engineering. There's some more nuanced and balanced primers on the topic (e.g. the Wooldridge book), but an awful lot more writing on AI Winters and what not.
This snowballs into a lack of awareness of the overlap between "AI" and, say, the body of knowledge in statistics.
In other words: Sure, I agree - HN collectively is less knowledge on this than on some other topics (myself included!), and it's worth taking stock of that, I suppose.
These silly generalisations about HN are pointless. HN has a pretty diverse set of worldviews, approaches, and levels of expertise across a range of topics. There is zero point in lumping people into a single, or even a majority, behaviour.
The idea that generalities can't be true or useful is pretty insidious, since it prevents reasoning about groups.
This idea was not mentioned, whether or not you attest that it was.
The onus [1] should go the other way.
____________
[1] Today's conversation-derailing trivium: "onus" means ass in Greek. As in donkey.
A colleague at my first job was working on some software that wrapped around other software, and he dubbed it "Burrito". When his project turned out not to be as useful as expected, he was pleasantly surprised that he could also use the original meaning of the word burrito, being "little donkey" in Spanish [1].
This means we have to be lucky to have some expert with current knowledge join in. Given that most people here are billionaires with a passion for dynamic typing, and not philosophers of mind, I am actually fairly surprised at the reasonable level of discussion.
Also, I doubt that referencing prior literature in the context of cognition and consciousness would be of much help. HN writes a lot of nonsense about it, but so have many professional philosophers. After sifting through all of that, one still has to add a lot of convincing arguments to advance an unpopular opinion.
A friend once suggested to build forum software that would allow discussions to reach a certain level of trustworthiness or truth. After several lengthy discussions I still doubt that such a thing is feasible. But it might spark your interest? :)
What does it mean to understand something? To think? How is a human any different than predicting the next “thing” given a series of inputs?
For a lot of us, LLM behavior and thinking are so plainly and obviously dissimilar that the endless comparisons look somewhere between naive or manipulative, depending on who’s making them.
If you turn off the power to the computer hosting the LLM, you can also be sure it doesn't think.
But what's obvious to you is not obvious to everyone.
Even though you don't have to mathematically prove a goldfish doesn't walk for me to believe you, that's because we can both agree on a very good definition of walking. If it came down to it, I am sure we could sit down with pen an paper and some textbooks and agree upon a physical definition as robust as any. We're just skipping that step because it's been done before by others.
I am positive that we can't agree on a robust definition of thinking, because the definition of thinking is beyond human understanding.
We have to keep an open mind. I don't believe it's likely that any LLMs are experiencing anything we'd call thought, but without knowing how thought works, it would be foolish of me to say it's impossible. The problem in AI discussions is not the positions people are taking but the certainty with which they are taking them.
In that case, shouldn't the people who say LLMs can think do the hard work of explaining what "thinking" means and why they er think it's done by LLMs?
Surely the default position should be that LLMs don't think because they don't belong to the class of entities that we know can think. If that default assumption is wrong, well, then, someone has to do the hard work of rejecting it. But just claiming that we think they think because who knows what thinking is, is just an excuse to not do the work.
We can be sure humans think. Does a crow think? We don't know. Does an LLM think? We don't know.
We can talk about more specific phenomena in contexts that demand it, if we have the data. But there is no way to say right now that an LLM does or does not think. All we can say is it seems unlikely.
Constantly moving the goal posts with this incessant "but what IS intelligence, though?!?" gets us nowhere. By your own argument we can never define intelligence, so we can never actually discuss it. LLMs consistently and constantly fall down on tasks that we would not expect a human to fail at, but you and people like you insist we cannot use this as evidence because you refuse to accept any defined terms.
I mean, remain as hopeful and uncritical as you want. But you're basically asking the rest of us "Who are you gonna believe, me or your lyin' eyes?" and that's gonna go about as well for you as it typically does.
I'm not talking about getting scammed by SBF or whatever. I'm saying that there doesn't exist a definition of "thinking" that we can use to include or exclude LLMs from the group "beings that think". It's counterproductive to say they do or don't and is a distraction from useful conversation.
To say I'm not being critical when I clearly said I don't think it's likely that LLMs think feels disingenuous and, frankly, uncritical.
> By your own argument we can never define intelligence
In no way does that extend from my argument. Intelligence and thinking are separate phenomena, and I didn't say "never".
There's no reason to presume we all think the same way; rather the opposite, given how many ways in which humans are already known to think unalike to each other.
An easy way to prove the two are different is 'What is the equivalent to the endocrine system for an LLM and why doesn't it get tired like a mind?'
The logical problems that a lot of people run into are twofold:
1) At a very basic level, we modeled these architectures on brains (and named them after our brains), so at a very very basic metaphorical level, you can say 'these things are acting like our brains act.' This ignores the complexity of the brain.
2) We like to think that inconveniences of our minds (like the need for sleep) are "problems" that we can engineer away, rather than intrinsic components of the process. People won't like hearing this, but there's no formal reason why 'the need to sleep' is not a necessary criteria for 'being conscious,' because as far as we know, the only conscious beings out there also sleep. It's just human nature to assume we can engineer the 'good' parts of a biological system while avoiding the 'bad' while still essentially replicating the system, but that's usually not the case.
Are you suggesting that the brain is the only implementation that can “think”? Or that an endocrine system is required?
The difference isn’t the interesting, or necessarily useful, part; the practical similarities of the output are. If we can make a system that appears to "think", I can’t understand how it can be so easily dismissed when we don’t know what’s going on in it or us.
What else does (please do not mention any deterministic counting machines like semiconductors - neurons are not deterministic and thought isn’t a deterministic set of calculations).
I would argue that the difference in process is extremely important, as evidenced by how easily we are fooled by optical illusions.
Just because we think two things appear the same does not mean they are, and understanding why they appear the same but are different relies on a study of the “how”
To put your argument differently - “if we can make drawings that appear to be three dimensional to our eyes, why bother understanding the difference between 2-D and 3-D space?”
I have to mention it, because it's physics, and related to the current implementation of LLMs: random numbers are possible, and used to break determinism. Intel CPUs use thermal noise to generate random numbers [1]. With silicon, randomness is a free choice, not an impossibility. LLMs front ends, like anything from OpenAI, use random numbers to get non-deterministic output, which is also the input of the next word, and the context of the next response, resulting in output that's not deterministic, with broad divergence [2]. Both systems are somewhat bound by the "sensibility"/logic of the output though, of course.
> as evidenced by how easily we are fooled by optical illusions
This isn't neccesarily unique to humans [3]. Do you have a specific illusion in mind? Many are related to active "baselining", and other time response things, that happens in our eyes, with others being an incorrect resolution of real ambiguity, from a two sensor system, that any sensor system will struggle with.
> Just because we think two things appear the same does not mean they are
I don't think anyone is suggesting they're the same, but I see many people suggesting, with seemingly undue confidence, that they're completely unrelated, which would require an understanding of either system that we don't have.
> To put your argument differently - ...
My argument is, if it's 3d to your eyes, then that understanding, and relation, between 2d and 3d space already exists in the system, to some degree.
[1] https://www.intel.com/content/www/us/en/developer/articles/g...
[2] https://www.coltsteele.com/tips/understanding-openai-s-tempe...
[3] https://blog.frontiersin.org/2018/04/26/artificial-intellige...
A dot product run on Intel CPUs are intended to always be the same no matter what. The heat signature stuff isn’t changing the way circuits do math.
As to optical illusions, the point I was making was that “the way humans practically perceive something” is not a sufficient way to measure the similarness of two things, since our perceptions are so often tricked (LLMs are literally trained specifically to trick you into thinking you’re talking to sentience).
I also don’t think they are completely unrelated at all. As I said I know we designed one to be like the other. It’s just by no means “the same.” AI may one day do convergent evolution toward how brains think, but it’s still important to recognize that’s convergent evolution.
> It’s just by no means “the same.”
The implementation is not the same. Everyone agrees with that. The concept being compared is "thought", not "implementation of thought". Maybe I'm lost.
I would argue expanding the definition to include the kind of things that semiconductors do really dilutes the meaning of "thinking" to be near meaningless.
Maybe to put it succinctly - no computer has ever done anything without a human input (even if it's millions of layers abstracted), but thinking just happens spontaneously.
If that's not a sufficient condition to differentiate 'thinking' from 'calculating,' then IDK what 'thinking' even means then.
It’s a common experience that candy tasted sweeter to you as a child, right?
How does that work in a transformer?
Indeed, but I'm not trying to with that comment, which is just about how my own seems to work on self-reflection. That said, I do have reason to doubt the accuracy of human introspection of our own thought processes, and therefore my own judgment in this may also be flawed.
Certainly, here's an exponential equation that can be challenging to solve for integer solutions:
2^x+3^y=7^z
This equation involves three variables, x, y, and z, and requires finding integer values for these variables that satisfy the equation. This type of equation is known as a Diophantine equation, and solving it can be quite challenging, especially for larger values of x, y, and z.
"
I mean, it's definitely not a Diophantine equation and solving it is definitely not challenging -- (2, 1, 1) happens to be an easy solution -- but I want to say it probably doesn't have any other solutions but I don't see a great way to prove it...
LLMs approach learning differently -- whatever reasoning they do possess is an emergent property that arises as they get better and better at language generation. In other words, unlike humans, LLMs learn to "speak" before they exhibit behaviors that look like logical reasoning. Humans do the opposite.
It would be both amusing and disturbing to see older Usenet, Slashdot, and Reddit conversations turned into rare and valuable resources.
He'd say: You are conscious. You are also a computer. You are not, as far as we can tell, conscious by virtue of being a computer. They're just two things that happen to be true about you but they are not linked in that super-direct way.
Similarly, you can think. And you can generate text. But you are not, as far as we can tell, thinking by virtue of generating text.
A lot of people think in concepts without words at all in many scenarios. E.g. when you cook an egg, do you reason (in words) with yourself to decide what temperature to set the dial on the stove? Or do you just do it “without really thinking”?
Does your brain respond with annoying disclaimers when asking yourself medical questions? Does your brain refuse to entertain an idea if it’s questionable?
If you're not with us you're...
If you scratch my back I'll..
A bird in the hand is worth...
What is a bird in the hand worth? How did you know that?
Your response to "not all problems are solved through generating text" is "but what if the problem is specifically to generate text?" Sure? Yes, if you ask me to complete a sentence, I complete a sentence. That doesn't indicate that human brains are reducible to text generators.
If anything, what's interesting about your example is that it might actually in some ways demonstrate the opposite of your point. How many instances are there of people memorizing text or song lyrics and singing along with an artist dozens of times, and only later sitting down and thinking about what the words actually mean? For humans, it's well understood that memorization and repetition of a text is not necessarily the same thing as understanding the concepts within that text.
I feel confident I could find you some kids that know that "a bird in the hand is worth two in the bush" that have never actually thought about what that phrase is intending to convey.
I think you've made my point quite well.
That humans are capable of memorizing text? Is that a thing anyone was debating?
Sure there's other stuff going on, but a lot of it is language and a lot of that subset seems to be quite closely approximated by these larger LLMs, warts and all.
If I told you that an LLM was basically a Markov chain because there are situations where the Markov chain and LLM produce similar output and where you could reasonably argue that it's possible the LLM is working the same way that the Markov chain is working -- you would (correctly) say I was oversimplifying what's going on in GPT. It would be a terrible comparison to make. Similarly, if you say that human reasoning is the result of an LLM, and what you're actually saying is that in a subset of situations humans produce output that could theoretically be working the same way as an LLM, I'm gonna say you're oversimplifying how human brains work.
Very clearly, there is more going on in a human brain than language, evidenced by the fact that our brains are larger than our language centers and we can literally measure which parts of the human brain experience the most activity when we're tackling different tasks.
> Sure there's other stuff going on, but a lot of it is language and a lot of that subset seems to be quite closely approximated by these larger LLMs, warts and all
If you define your tests and scenarios to specifically encompass only the situations in which similar outputs are produced, then sure. But that's not a very strong argument for you to make. It's like saying chickens are basically the same as fish since both of them lay eggs. The phrase "sure there's other stuff" is doing a lot of work there, because "there's other stuff" is what we're all saying when we say that LLMs aren't just primitive humans -- LLMs are different from humans in the sense that when we reason, there's other stuff going on and the entire process is not reducible to only language generation.
----
Of course, absent from this conversation is the fact that the way our language center develops is different from how LLMs are trained and even in the parallels you're drawing, LLMs demonstrate different strengths and weaknesses from humans -- humans demonstrate reasoning capabilities faster than proficiency with language/text, LLMs demonstrate proficiency with text faster than they demonstrate reasoning capabilities. Even in situations where both LLMs and humans predict text, it's likely that we're using different strategies to do so given our differing capabilities.
Look I am not even making a claim about whether GPT can reason. Defining intelligence through a purely human lens would be unimaginative and needlessly narrow. Whether GPT actually reasons or just appears to is a different conversation. I haven't touched that conversation, all that I'm saying is, it's very obvious that whatever GPT is doing, it is different from how human brains work.
I was pushing back against something that might not have been there initially: an unwillingness to accommodate the idea that a fairly significant part of our intelligence is actually stored in our language and that it can seem subjectively (perhaps excluding those with no internal monologue) that we recall that knowledge in a manner which is close enough to the one we have managed to replicate in the larger LLMs.
This part of our function could well be a lazy optimization that in reality sits on top of our reasoning capabilities to save energy, but the point of the trite task with the bird in the hand was just to demonstrate that it seems to play a fairly significant role. I'm out of my depth entirely with respect to the actual form and function of the brain as per the state of research today.
I'm willing to admit though that when I replied to you I was probably really replying to a lot of other commenters, many of whom had stronger objections to the idea that LLMs are now knocking on the door of intelligence at least to the point where we feel the need to redefine it.
Many scenarios, not all scenarios.
If you say “choose X or Y and then give three supporting arguments” the LLM does not write out the supporting arguments, but the vector space determining the initial one word answer of X or Y does include the embedded awareness of those arguments to different levels of specificity and relevance.
The annoying disclaimers and refusals are not fundamental to LLMs but specific LLM services.
It's certainly possible that they're wrong, but my suspicion is that at the very least the form their internal monologue is taking is not one of consciously simulating hypothetical conversations in their head.
Additionally, we know that human decisions can in some situations be influenced by stimuli that take effect before the brain has even consciously registered that a decision needs to be made. It seems pretty safe to say that those stimuli are not being processed using a language model.
Could you expand on that? I don’t see how it follows. If there is unconscious output, why wouldn’t it suggest there’s something analogous to a “unthinking” language model stuck behind our conscious, self reflective, bits? How does unconscious output prove that it’s not something similar to a language model?
Or, are you speaking on a technicality, saying it’s not a literal language model? I think the idea is something similar and, obviously, multi modal, not a literal LLM/text predictor implementation like runs on GPUs.
----
So on one hand, you could call that a technicality, on the other hand I would say that if we reduce "like an LLM" to mean "responds to a stimuli by generating a response", at that point "like an LLM" is so broad as to be meaningless. Sure, we can define "humans use a language model" as "human brains generate signals to other parts of the human brain in response to stimuli" -- but that's not really a useful thing to say when we're comparing human reasoning to GPT-4.
If we're talking so broadly about reasoning, then it would be just as accurate to say that human brains are like Markov chains. After all, Markov chains interpret inputs and transform them into other signals that they send to other parts of the chain. But if I told you that GPT-4 was just a Markov chain you would (rightfully) disagree with me on that and you would (rightfully) say that I was oversimplifying what's going on within GPT-4's model.
Typically, when people say that human beings are text generators, they're talking in the context of GPT-4; they're trying to claim that GPT-4's behavior is analogous to human behavior. Whether or not GPT-4 can reason under any definition of reasoning (direct analogy to human thought should not be the only way we think about reasoning), the way GPT works and the way that it's trained and taught is extremely alien to how human brains work and how humans learn. Broadening the definition doesn't change the fact that LLMs and human brains are very different in practical ways that influence their observable behaviors.
----
I don't think people are consciously trying to do this, but the debate over whether humans are LLMs often ends up feeling like an attempt to simultaneously broaden and narrow a definition at the same time. What people seem to want to do is to describe LLMs extremely broadly but then assume that conclusions they draw from that broad definition are necessarily applicable to a very narrow definition of an LLM (specifically, GPT and similar models). But once we define an LLM broadly enough to encompass the entirety of human reasoning, we have also broadened the category enough to encompass a lot of processes that no one would call reasoning. What doesn't transform inputs into other forms and pass them along to another object? An AC/DC electrical converter does that, but that doesn't mean it can think.
It's like claiming, "computers are made of carbon and so are humans, so computers are basically the same as humans." Well, wait a second -- the category as you've defined it is so broad that objects within that category can no longer be assumed to share other attributes with each other.
> "like an LLM" is so broad as to be meaningless.
I don't think this is a fair interpretation, in the context of this comment chain. To me, a charitable interpretation is that "like an LLM" means no feedback loop, no refinement, just input to output, like an LLM.
I point my brain at problems, using my conscious intent, and I see answers. I don't really consider myself consciously involved in those answers, and I do have to check them, and sometimes iterate, with that iteration almost always involving some solidification of the answers/concepts by externalizing them in writing, drawing, etc. Point at solidification, get response, repeat. The "me" in my head is the one talking, pointing, receiving, checking, and mixing. The "here's your answer" intuition that I experience is the "like an LLM thing". My understanding of my own thought process follows the historic interpretation: the pre-frontal cortex is a new attachment, and probably the one that is the active participant, and planner, that is "me". The fast, but unrefined, input->output system being the rest of it, and is probably closer to the experience of a monkeys day to day activities.
But that's not what an LLM is, an LLM is a large language model. The distinguishing characteristics between an LLM and other AIs are not whether or not an LLM has a feedback loop. Lots of AI categories (arguably the majority of AIs) generate output from a single set of inputs as a single "operation" (ignoring the fact that most neural networks including LLMs internally have multiple layers and do in fact have multiple transformation steps, but whatever, we can treat that as one step for the purposes of conversation).
If anything, LLMs are the exception here in that they very often do have a feedback loop during normal usage; they are most commonly used in a conversational context where they generate the next "chunk" of a conversation after being fed back their previous answers alongside the followup responses of the person working with them. Arguably the feedback loop of an LLM is that as a conversation progresses, its future output is based on its previous output, which very literally becomes its new input after being marked up and extended by a human being. That is notably more feedback than many other AIs get.
Doubly so if you're trying to argue that an LLM can encompass a multi-modal setup, because suddenly refinement and feedback between multiple models is a core part of the final product.
So I'm not sure I agree that an LLM fits into that category in the first place, but even assuming it does, if your definition of an LLM is is just that it takes an input and turns it into an output as a single step without help... that is just such a broad category, I would guess the majority of AIs fall into that category. It's not what makes LLMs special; most predictive neural networks take a single set of inputs and generate a single set of outputs without intermediary human input. What makes LLMs interesting as a category is the training and structure and quirks of how they work, what makes them interesting is the differences between LLMs and other AI techniques.
To jump back to the Markov chain again, Markov chains do not have a feedback loop or refinement step: they take an input and map it to an output as a single step. Are Markov chains LLMs? Is any data transformer that operates in a single step an LLM? That is a really broad definition to use.
----
And we run into the same problems, because if you're arguing that a human brain is like an LLM and what you're really saying is, "it's kind of like an AI in general" -- well, that doesn't really say anything about whether a human brain is specifically similar to a system like GPT. There are lots of different ways to build a neural network, LLMs are one strategy.
When you say this:
> The "here's your answer" intuition that I experience is the "like an LLM thing".
What this sentence is actually saying is that you experience conclusions and knowledge where you don't know the source or where the process of your brain generating a decision or information happens unconsciously. The sentence is just saying that there are unconscious parts of your brain where you are not an active participant in the thinking process and it feels like the a spontaneous transformation from input to output.
But does that really strike you as a strong indicator that your brain is like an LLM, or does that sound more like a general description of almost any black-box system where the mechanisms are hidden and you can only examine the inputs and outputs? To say that there are parts of our thinking process that we're not able to consciously observe or examine is really just saying that parts of our thinking process resemble a black-box oracle. It's not making a strong claim about GPT.
But we rarely do, outside of (maybe) therapy.
“Can we agree that of the many different things folks mean when they say your consciousness is ‘causally generated by your brain,’ that at least we can say your brain in its present state is only capable of generating exactly one conscious process at a time, namely you?”
So it's a sort of causal sufficiency, Searle believes that there is something about this squishy machine that the physics is doing that makes it conscious, and that's being caused very directly by the squishy machine. It's not some magic.
On the other hand, if you believed in souls, you might think that a given brain contained both the soul of a human person, and the soul of some demon possessing them—two conscious processes, one of which was stuck “along for the ride.” [Searle is also interested in the cases of split-brain patients where you have a similar phenomenon because essentially two separate brains inhabit a single body. He has I think mentioned at a talk that one of the interesting things about consciousness to him is that it all gets Unified, so it's interesting that if the two parts of the brain can talk to each other they merge their consciousnesses into one more powerful consciousness, rather like (my analogy, not his) how if you have two water droplets on a plastic plate and you push them with a toothpick together, at some point they merge into one bigger drop.]
Now the Chinese room thought experiment is not about the computational model of consciousness—not directly! It was always phrased as a rebuttal of the Turing test in particular, and the computational model only indirectly after the Turing test falls. Note that the Turing test has no direct beef with the casual sufficiency axiom, which is why it went unstated originally. According to the Turing test, written text goes in, written text comes back out, a dialogue appears to happen to the outsider, this is sufficient to conclude conscious understanding of the language used, which confirms consciousness.
Searle’s objection is, “if it's really all about inputs and outputs and not how I get it done, then you've left out what for me feels like the most important part about understanding a language: understanding it is part of how I get it done!” Right? You understand English because you can phrase your ideas into it, you can mold it to suit you, and it can (when heard) so mold you—it’s not just because the words can come out of your mouth, triggered by other words that came in through your ears. The Turing test has always ever stated “don't worry what happens in the middle” and Searle is saying “but for me understanding is part of the details that are happening in the middle!”
So where does the computational theory come in? It comes in because Searle wants to make this argument rigorous! He says, “if your computational theory is true, then there is in fact another way that I can speak a language—words in, words out—where I don't understand a word of Chinese. I am Turing complete, I could memorize a program for speaking Chinese and happen to execute it flawlessly and at no point would my ideas get into Chinese, at no point would the things that I heard mold me. So the Turing test is crap!”
The prominence of the words “I, me” is what invokes the idea of causal sufficiency here, “I have the right inputs and outputs, but I don't understand.” The Turing test is not able to distinguish between multiple consciousnesses instantiated by the same hardware, if such is even possible. The inputs and outputs go into the same box, as far as Turing is concerned as long as only one conversation happens, there's only one person in there.
This does do a great job of defanging the Turing test, because all of the ways that you might weaken the concept to address this major limitation do make it sound completely tautological. “Yeah well something consciousnessy is happening in that box but I don't know what.” / “Okay then why are we even talking about it.” / “Because computers can speak!” / “Right, so we care that computers can speak because they can speak?” / “No, like, we gotta give them rights now, or some shit.” / “John Searle already has human rights, if he's memorized a program that lets him speak Chinese without understanding it, you're saying we need to give that program human rights?” / “Yeah!” / “So uh is it murder if Searle decides that running the program is no longer fun? Is he a slave to this program forever?” / “Uhhhh...”
It doesn't defang the computational model, not directly. But the computational model does imply that VMs exist. We use them every day! And that's all that the Chinese room is, it's running a VM inside of another computer, one consciousness carrying a separate consciousness inside of it, a willful sort of demonic possession. The only thing the Chinese room has to say about this, is that we don't use our language very well if it is true. Philosophers who believe in that will need to generate an alternative language that is able to distinguish between “I am doing it” and “I am sustaining a daemon who is doing it,” because for them that's a real difference, you might have a hundred consciousnesses in your head that you don't have direct access to. That is a necessary part of believing that consciousness is software, you don't know if you're in a VM inside your brain, you don't know if something you're doing is actually secretly a Brainfuck program instantiating another VM inside your consciousness, software embeds within software, that's a core feature of software.
But of course Searle thinks that that's kind of ridiculous because he thinks that it's obvious that consciousness is something that the squishy wetware of the brain does, and this forces him to believe in that causal sufficiency—“my brain is only sustaining one consciousness, namely me,”—which the computationalists cannot ever agree with because that's not how software works. Anyone who believes in that causal sufficiency, even if they don't have the same basis that Searle does for believing in it, also thinks that the computationalists are ridiculous.
But the point is, that's happening at a level way above the Chinese room argument, Chinese room baseball came and went, now this is a whole separate game being played at the same ballpark afterwards.
That said, this all seems like a 'fallacy of composition'. Humans are not LLMs, so much should be obvious, but at the same time concluding that they are completely different just feels wrong. The mistakes LLMs make feel very similar to what humans do when they don't have the time to deeply think about a problem and just give you the best guess that pops into their head. Humans will have other systems on top that allow them deeper reasoning, but the language generation really doesn't feel all that different from what LLMs do.
That aside, humans interact with the world, they get instant feedback on what of their predictions is right or wrong. LLMs are stuck with just static training data that might simply not be enough to develop higher level reasoning skills.