The LLMentalist Effect (2023)
softwarecrisis.dev
softwarecrisis.dev
> 2) The intelligence illusion is in the mind of the user and not in the LLM itself.
I've felt as though there is something in between. Maybe:
3) The tech industry invented the initial stages a kind of mind that, though misses the mark, is approaching something not too dissimilar to how an aspect of human intelligence works.
> By using validation statements, … the chatbot and the psychic both give the impression of being able to make extremely specific answers, but those answers are in fact statistically generic.
"Mr. Geller, can you write some Python code for me to convert a 1-bit .bmp file to a hexadecimal string?"
Sorry, even if you think the underlying mechanisms have some sort of analog there's real value in LLM's, not so psychics doing "cold readings".
I do think there is a degree of mentalist-like behavior that happens, maybe especially because of the RLHF step, where the LLM is encouraged to respond in ways that seem more truthful or compelling than is justified by its ability. We appreciate the LLM bestowing confidence on us, and rank an answer more highly if it gives us that confidence... not unlike the person who goes to a spiritualist wanting to receive comforting news of a loved one who has passed. It's an important attribute of LLMs to be aware of, but not the complete explanation the author is looking for.
Outside of grammar, you can hear a lot of this when people talk; their sentences wander, and they don't always seem to know ahead of time where their sentences will end up. They start anchored to a thought, and seem to hope that the correct words end up falling into place.
Now, does all thought work like this? Definitely not, and more importantly, there are many other facets of thought which are not present in LLMs. When someone has wandered badly when trying to get a sentence out, they are often also able to introspect and see that they failed to articulate their thought. They can also slow down their speaking, or pause, and plan out ahead of time; in effect, using this same introspection to prevent themselves from speaking poorly in the first place. Of course there's also memory, consciousness, and all sorts of other facets of intelligence.
What I'm on the fence about is whether this point, or your point actually detracts from the author's argument.
I can create a mad-libs program which dynamically reassembles stories involving a kind and compassionate Santa Claus, but that does not mean the program shares those qualities. I have not digitally reified the spirit of Christmas, not even if excited human kids contribute some of the words that shape its direction and clap with glee.
P.S.: This "LLM just makes document bigger" framing is also very useful understanding how prompt injection and hallucinations are constant core behaviors, which we just ignore except when they inconvenience us The assistant-bot in the story can be twisted or vanish so abruptly because it's just something in a digital daydream.
So you simultaneously exist and don’t exist. Sorry about this, your post took me on this tangent.
The real question that gets billions invested is "Is it useful?".
If the "con artist" solves my problem, that's fine by me. It is like having a mentalist tell me "I see that you are having a leaky faucet and I see your future in a hardware store buying a 25mm gasket and teflon tape...". In the end, I will have my leak fixed and that's what I wanted, who care how it got to it?
The graveyards of startups are littered with economically infeasible solutions to problems.
I wouldn't read too much into customers knowing what they want.
https://www.afstores.com/the-intriguing-story-of-how-the-sho...
or if those sources aren't your cup of tea, how about Fox News
https://www.foxnews.com/lifestyle/meet-american-invented-sho...
> The public — naturally — hated the idea.
The amount of money being invested is very clearly disproportionate to the current utility, and much of it is obviously based on the premise that LLMs can (or will soon) be able to think.
I don't think this is so clear. At least, if by "current utility" you also include potential future utility even without any advance in the underlying models.
Most of the money invested in the 2000 bubble was lost, but that didn't mean the utility of the internet was overblown.
This seems to be a hard assumption the entire post, and many other similar ones, rely upon. But how do you know how people think or reason? How do you know human intelligence is not an illusion? Decades of research were unable to answer this. Now when LLMs are everywhere, suddenly everybody is an expert in human thinking with extremely strong opinions. To my vague intuition (based on understanding of how LLMs work) it's absolutely obvious they do share at least some fundamental mechanisms, regardless of vast low-level architecture/training differences. The entire discussion on whether it's real intelligence or not is based on ill-defined terms like "intelligence", so we can keep going in circles with it.
By the way, OpenAI does nothing of this, see [1]:
>artificial general intelligence (AGI)—by which we mean highly autonomous systems that outperform humans at most economically valuable work
Neither do others. So the author describes "tech industry" unknown to me.
From the article:
> The field of AI research has a reputation for disregarding the value of other fields [...] It’s likely that, being unaware of much of the research in psychology on cognitive biases or how a psychic’s con works, they stumbled into a mechanism and made chatbots that fooled many of the chatbot makers themselves.
Just because cognitive scientists don't know everything about how intelligence works (or on what it is) doesn't mean that they know nothing. There has been a lot of progress in cognitive science, in the last decade in particular on reasoning.
> based on ill-defined terms like "intelligence".
The whole discussion is about "artificial intelligence". Arguably AI researchers ought to have a fairly well defined stance of what "intelligence" means and can't use a trick like "nobody knows what intelligence is" to escape criticism.
There is a lot of work to define intelligence, both in the context of cognitive science and in the context of AI [0].
I haven't spent enough time looking for a good review article, but for example, this article [1] says this: "Intelligence in the strict sense is the ability to know with conscience. Knowing with conscience implies awareness." and contrasts it with this: "Intelligence, in a broad sense, is the ability to process information. This can be applied to plants, machines, cells, etc. It does not imply knowledge." If you are interested in the topic the whole article is interesting and worth reading.
[0] https://scholar.google.com/scholar?hl=en&as_sdt=0%2C5&as_ylo...
[1] https://www.cell.com/heliyon/fulltext/S2405-8440(21)00373-X
It's what makes most sense to me.
> Intelligence in the strict sense is the ability to know with conscience. Knowing with conscience implies awareness.
This one I would disagree with.
> Intelligence, in a broad sense, is the ability to process information. This can be applied to plants, machines, cells, etc. It does not imply knowledge.
This one starts with one part of what I consider intelligence, but besides processing it also needs to be able to use the information to solve problems. Because you could process information, which everything in the World actually does technically, but you would not be using that information to do anything.
So ultimately, maybe it would be best to define it as it's the ability to take in and use information to solve some sort of problem.
Author's claim is pretty strong: that human intelligence and what is called GenAI have no common ground at all. This is untrue at least intuitively for the entire field of ML. "Intuitively" because it cannot be proven or disproven until you know exactly what is human intelligence, or whatever the author means by thinking or reasoning.
>Arguably AI researchers ought to have a fairly well defined stance of what "intelligence" means and can't use a trick like "nobody knows what intelligence is" to escape criticism.
If you don't formalize your definition, the discussion can be easily escaped by either side by simply moving along the abstraction tree. The tiresome stochastic parrot/token predictor argument is about trees while the human intelligence is discussed in terms of a forest. And if you do formalize it, it's possible to discover that human intelligence is not what it seems either. I'm not even starting on the difference between individual intelligence, collective intelligence, and biological evolution, it's not easy to define where one ends and another begins.
AI researchers mainly focus on usefulness (see the definition above). The proof is in the pudding. Philosophical discussions are fine to have, but pretty meaningless at best and designed to support a narrative at worst.
I feel that you're being unfair to the author here; the quote you responded to in your GP post alluded to "reason or think", and their argument is that LLMs don't. This is more specific than the sweeping statement you attribute to them.
> AI researchers mainly focus on usefulness (see the definition above).
And usefulness is something the article doesn't touch on, I think? The point of the article is that some users attribute capabilities to LLMs that can be explained with the same mechanisms as the capabilities they attribute to psychics (which are well understood to be nonexistent).
> Philosophical discussions are fine to have, but pretty meaningless at best and designed to support a narrative at worst.
Why this is relevant to AI research is (in my interpretation) that it is known to be hard for humans collectively to evaluate how intelligent an entity really is. We are easily fooled into seeing "intelligence" where it is not. This is something that cognitive scientists have spent a lot of time thinking about, and may perhaps be able to comment on.
For what it's worth I've seen cognitive scientists use AI to do cools stuff. I remember seeing someone who was using AI to show that it is possible to build inference without language (I am speaking from memory here, it was a while ago and I lost the reference sadly so I hope I'm not deforming it too much). She was certainly not claiming that her experiments showed how intelligence worked, but only that they pointed to the fact that language does not have to be a prerequisite for building inferences. Interesting stuff though without the sensationalism that sometimes accompanies the advocacy of LLMs.
Is the reputation warranted? Just a US thing? Or maybe the question is "since when did this change"? because in the mid 2000s in france at least, llm research was led by cognitive psychology professors who dabble in programming or had partnerships with a nearby technical university.
My experience isn't much, since I am neither doing AI nor cognitive science but I have seen cognitive scientists do cool stuff with AI as a means to study cognitive science, and I have seen CS researchers involving themselves into the world of AI with varying success.
I would not be as emphatic as the author of the article but I would say that a good portion of AI research lost the focus on what is intelligence and instead just aimed at getting computers to perform various tasks. Which is completely fine until these researchers start claiming that they have produced intelligent machines (a manifestation of the Dunning Kruger effect).
https://arcprize.org/blog/openai-o1-results-arc-prize
I wouldn't say it's full agi or anything yet, but these things can definitely think in a very broad sense of the word
People tend to do so all the time, with games for example.
I'm talking transfer learning and generalization. A human who has never seen the problem set can be told the rules of the problem domain and then get 85+% on the rest. o3 high compute requires 300 examples using SFT to perform similarly. An impressive feat, but obviously not enough to just give an agent instructions and let it go. 300 examples for human level performance on the specific task, but that's still impressive compared to SOTA 2 years ago. It will be interesting to see performance on ARC-AGI-2.
You should try to play chess yourself, and then tell me you think these things aren't intelligent.
I don't mean anything by the following either, other than, the goalposts have moved:
- This doesn't say anything about generalization, nor does it claim to.
- The occurrences of the prefix general* refer to "Can fine-tuning with synthetic logical reasoning tasks improve the general abilities of LLMs?"
- This specific suggestion was accomplished publicly to some acclaim in September
- To wit, the benchmark the article is centered around hasn't been updated since since September, because the preview of the large model accomplishing that blew it out of the water, 33% on all at the time, 71%: https://huggingface.co/spaces/allenai/ZebraLogic
- these aren't supposed to be easy, they're constraint satisfaction problems, which they point out are used on the LSAT
- The major other form of this argument is the Apple paper, which shows a 5 point drop from 87% to 82% on a home-cooked model
If the model has in its internal world model knowledge it likely does not know how to solve a coding question, but the RLHF stage has reviewers rate refusals lower, it would in turn force its hand when it comes to tricks it knows it can pull based on its model of human reviewers. It can only implement the surface level boilerplate and pass that off as a solution, write its code in APL to obfuscate its lack of understanding, or keep misinterpreting the problem into a simpler one.
A psychic that read on ten thousand biographies might start to recall them, or he might interpolate the blanks with a generous dose of BS, or more likely do both in equal measure.
Isn't that usually by not even trying, and delegating the work regular programs?
An LLM is at best, a possible future component of the speculative future being sold today.
How might future generations visualize this? I'm imagining some ancient Greeks, who have invented an inefficient reciprocating pump, which they declare is a heart and that means they've basically built a person. (At the time, many believed the brain was just there to cool the blood.) Look! The fluid being pumped can move a lever: It's waving to us.
Before intuitive computing, the best we could do with word problems was Wolfram-esque regex stuff, which I’m guessing we all know was quite error-prone. Now, we have agents that can take quite vague word problems and use any sequence of KB/web searches, python programs, and further intuitive reasoning steps to arrive at the requested answer. That’s pretty impressive, and I don’t think “well technically it relies on tools” makes it less impressive! Something that wasn’t possible yesterday is possible today; that alone matters.
Re:general skepticism, I’ve given up on convincing people that AGI is close, so all ill say is “hedge your bets” ;)
if you wish for GP to do that, ask them to do that
There is nothing of substance in this and it feels like the author has a grudge against LLMs.
You can spot them easily, because instead of critiquing some specific thing and sticking to it, they can't resist throwing in "obviously, LLMs are all 100% useless and anyone who says otherwise is a Tech Bro" somewhere. Like:
completely unknown processes that have no parallel in the biological world.
c'mon... Anyone who knows a tiny bit about ML knows that both of those claims are just absurdly off base.Agreed on the experiments. What would they look like? Can a chat bot give the same info without any bedside manner?
Edit: Don't conflate mechanisms with capabilities.
To take your analogy even further, it is like asking when is the plane going to improve enough that it can really fly by flapping it's wings.
> 2) The intelligence illusion is in the mind of the user and not in the LLM itself.
3) The intelligence of the users is illusion either?
Someone should write a blog post about this to warn humanity.
at that point, capital can tell most humans to just go away and die, and can use their technology to protect themselves in the meantime
What LLMs seem to emulate surprisingly well is something like a person's internal monologue, which is part of but not the whole of our mind.
It's as if it has the ability to talk to itself extremely quickly and while plugged directly into ~all of the written information humanity has ever produced, and what we see is the output of that hidden, verbally-reasoned conversation.
Something like that could be called intelligent, in terms of its ability to manipulate symbols and rearrange information, without having even a flicker of awareness, and entirely lacking the ability to synthesise new knowledge based on an intuitive or systemic understanding of a domain, as opposed to a complete verbal description of said domain.
Or to put it another way - it can be intelligent in terms of its utility, without possessing even an ounce of conscious awareness or understanding.
That settled this question for me.
Language, as a problem, doesn’t have a discrete solution like the question of whether a list is sorted or not.
Seems weird to compare one to the other, unless I’m misunderstanding something.
What’s more, the entire notion of a sorted list was provided to the LLM by how you organized your training data.
I don’t know the details of your experiment, but did you note whether the lists were sorted ascended or descended?
Did you compare which kind of sorting was most common in the output and in the training set?
Your bias might have snuck in without you knowing.
It is nothing new and has been well established in the literature since the 90s.
The shared article really is not worth the read and mostly uncovers an author who does not know what he write about.
LLMs didn’t exist in then. Attention only came out in 2017…
The network itself can be trained to solve most functions (or all, I forget precisely if NNs can solve all functions)
But the language model is not necessarily capable of solving all functions, because it was already trained on language.
If every pair of digits appears sorted in the dataset, then that could still be “just” a stochastic parrot.
I’m kind of interested to see if an LLM can sort when the dataset specifically omits comparisons between certain pairs of numbers.
Also I don’t think OC was responding to commenters, but the article
But by specifically avoiding certain cases, wet could verify if the model is generalizing or not.
As for avoiding certain cases, that could be done to some extent. But remember that the untrained transformer has no preconception of numbers or ordering (it doesn't use the hardware ALU or integer data type) so there has to be enough data in the training set to learn 0<1<2<3<4<5<6, etc.
This is the kind of thing I’d want it to generalize.
If I avoid having 2 and 6 in the same unsorted list in the training set, will sets containing those numbers be correctly sorted in the same list in the test set and at the same rate as other lists.
My intuition is that, yes, it would. But it’d be nice to see and would be a clear demonstration of the ability to generalize at all.
For this hypothesis: The intelligence illusion is in the mind of the user and not in the LLM itself.
And yes, the notion was provided by the training data. It indeed had to learn that notion from the data, rather than parrot memorized lists or excerpts from the training set, because the problem space is too vast and the training set too small to brute force it.
The output lists were sorted in ascending order, the same way that I generated them for the training data. The sortedness is directly verifiable without me reading between the lines to infer something that isn't really there.
“the initial stages a completely new kind of mind, based on completely unknown principles, using completely unknown processes that have no parallel in the biological world.”
We just call it a neural network because we wanted to confuse biology with math for the hell of it?
“There is no reason to believe that it thinks or reasons—indeed, every AI researcher and vendor to date has repeatedly emphasised that these models don’t think.”
I mean just look at the Nobel Prize winners for counter examples to all of this https://www.cnn.com/2024/10/08/science/nobel-prize-physics-h...
I don’t understand the denialism behind replicating minds and thoughts with technology - that had been the entire point from the start.
But I don't see any discussion of multilayer perceptrons or multi-head attention.
Instead, the rest of the article is just saying "it's a con" with a lot of words.
LLMs write code, today, that works. They solve hard PhD level questions, today.
There is no trick. If anything, it's clear they haven't found a trick and are mostly brute forcing the intelligence they have. They're using unbelievable amounts of compute and are getting close to human level. Clearly humans still have some tricks that LLMs dont have yet, but that doesn't diminish what they can objectively do.
Different people have different definitions of intelligence. Mine doesn't require thinking or any kind of sentience so I can consider LLMs to be intelligent simply because they provide intelligent seeming answers to questions.
If you have a different definition, then of course you will disagree.
It's not rocket science. Just agree on a definition beforehand.
The machanism of intelligence is not understood. There isn't even a rigorous definition of what intelligence is. "All it does is combine parts it has seen in its training set to give an answer", well then the magic lies in how it knows what parts to combine, if one wants to go with this argument. Also conveniently, the fact that we have millions of years of evolution behind us, plus exabytes of training data over the years in form of different stimuli since birth gets shoved under the rug. I don't want to say that the conclusion is necessarily wrong, but the argument is always bad. I know it is hard to come to terms with the thought that intelligence may be more fundamental in nature and not exclusively a capability of carbon based life forms.
If intelligence were an objective property of the universe, we’d define it like mass or charge—quantifiable, invariant, fundamental. Instead, it shifts to match whatever we decide to measure. The instruments don’t quantify intelligence; they create it.
Yes, that is how LLM's work. They are trained with feedback loops to answer plausibly.
Can you please explain that a bit further? I don’t catch the connection you’re making between the conflation and being european.
> The LLMentalist Effect: how chat-based Large Language Models replicate the mechanisms of a psychic's con
> LLMs are a mathematical model of language tokens. You give a LLM text, and it will give you a mathematically plausible response to that text.
> The tech industry has accidentally invented the initial stages a completely new kind of mind, based on completely unknown principles, using completely unknown processes that have no parallel in the biological world.
Or maybe our mind is based on a bunch of mathematical tricks too.
But couldn't it be overfitting? LLMs are very good at deriving patterns, many of which humans simply can't tell apart from noise. With a few billion parameters and whatever black magic is going on inside CoT, it's not unreasonable to think even small amounts of fine-tuning combined with many epochs of training would be enough for it to conjure a compressed representation of that problem type.
Without an extensive audit, I'd be skeptical of OpenAI's claims, especially given how o1 is often wrong on much more trivial compositional questions.
What defines intelligence is generalization, the ability to learn new tasks from few examples, and while LLMs have made some significant progress here, they are still many orders below a child and arguably even many animals.
We can say that they're not "intelligent" because they're not capable of solving problems they can't map to something in their training at all, but that would also put 99.9% of humanity in the unintelligent bucket.
A human takes 14+ years until it's intelligent, also requires extensive training.
Some people used to push the theory that quantum probability was where free will and the soul reside. That is to say, people will imagine how the hard questions of old neatly fit into the hard questions of today. Nothing won't with that, it's how we explore different paths and make progress. But I'm not one of those exploring experts, so I'll wait for stricter definitions and experimental data.
It is not impossible I think, just require so much effort, talents, and funding that the last thing resembling such an endeavor was the Manhattan project. But if it succeeded, the impact could rival or even exceed what nuclear power had done.
Or am I deluded and there is some sort of fundamental limit or restriction on the transformer that would completely prevent this from the start?
Apart from that, I'm afraid that at this point research on sensory input apart from audio and visual needs much more advancement. For example, it's not clear to me what kind of data structure would be a good fit for olfactory or sensory training data
Touch and such can have some approximation done through various sensors like temperature, force, humidity, electromagnetic, etc.
Don't get me wrong I would be curious to see such research done to see whether it would improve anything above the stochastic parrot level - it's just going to take a while to figure out what is even relevant
But an LLM has no problem at all deciphering and processing and most importantly, responding meaningfully to all the ways we can use or encounter the word "lie". I contend that if a model large enough is trained on enough data, the concepts will automatically blend and explain each other sufficiently, or at least enough to cover your example and those similar to it.
Taste and olfactory are matters of chemical compositions. It will take an incredible effort but something similar to a mass spectrometer can be used to detect every taste and smell we can think of and beyond. How fast and how efficient they can be is probably the main challenge.
Touch is difficult. We don't even know fully why or how does an itch "work". But force, temperature, atmospheric, humidity sensors, etc are widely available. They can provide a crude approximation, imo.
Just off the top of my head. I am sure smarter people can come up with much more suitable ways to "embody" a machine learning model.
If reproducing the artifacts and failure modes of human modes of interpretation of this physical data (say, yanny/laurel, or optical illusions, or persistence of vision phenomena) is deemed important, that's another matter. If all that's required is a black-box understanding that is idiosyncratic to LLMs in particular, but where it's functionally good enough to be used as sight and hearing, then I don't see see why it can't be called "solved" for most intents and purposes in six months' time.
I guess it boils down to this: do you want "sight" to mean "machine" sight or "human" sight. The latter is a hard problem, but I'd prefer to let machines be machines. It's less work, and gives us a brand-new cognitive lens to analyse what we observe, a truly alien perspective that might prove useful.
No matter how you build it, it is still experiencing everything a human can experience. There's just no guarantee it would react the same way to the same stimuli. It would react in its own idiosyncratic way that might both overlap and contrast with a human experience.
A more "human" experience simulator would paradoxically be more and less authentic at the same time - more authentic in showing a human-style reaction, but at the cost of erasure of model's own emergent ones.
The reason we are doing all that is for its potential uses. Write letters, code, help customers, find information, etc... Even AGI is not about making artificial humans, it is about solving general problems (that's the "G").
And even if we could make artificial humans, there would be a philosophical problem. Since the idea is to make these AIs work for us, if we make these AI as human-like as possible, isn't it slavery? It is like making artificial meat but insist on making the meat-making machine conscious so that it can feel being slaughtered.
So instead of training it that way, the network can potentially be trained to "perceive" or "model" the reality beyond the digital world. The only way we know or have enough experience and data to do so is through our own experience. An embodied AI is what I think is required for anything to actually grasp the real concepts, or at least as close as possible to them.
And without that inherent understanding, no matter how useful a model is, it will never be a "general" inteligence.
But it doesn't have to be modeled after humans. The purpose of humans if we can call it that is to make more of itself, like all forms of life. That's not what we build robots for. We don't even give robots the physical abilities to do that. Giving them a human mind (assuming we could) would not be adequate. Wrong body, wrong purpose.
Stopped reading here. What is the mechanism in humans that enables intelligence? You don't know? Didn't think so. So how do you know LLMs don't have the required mechanism?
Not saying the author is wrong in general, but this kind of argument always annoys me. It's effectively a Forer statement for the "sceptics" side: It appears like a full-on refutation, but really says very little. It also evokes certain associations which are plain incorrect.
LLMs are functions that return a probability distribution of the next word given the previous words; this distribution is derived from the training data. That much is true. But this does not tell anything about how the derivation and probability generation processes actually work or how simple or complex they are.
What it does however, is evoke two implicit assumptions without justifying them:
1) LLMs fundamentally cannot have humanlike intelligence, because humans are qualitatively different: An LLM is a mathematical model and a human is, well, a human.
Sounds reasonable until you have a look at the human brain and find that human consciousness and thought too could be represented as nothing more than interactions between neurons. At which point, it gets metaphysical...
2) It implies that because LLMs are "statistical models", they are essentially slightly improved Markov chains. So if an LLM predicts the next word, it would essentially just look up where the previous words appeared in its training data most often and then return the next word from there.
That's not how LLMs work at all. For starters, the most extensive Markov chains have a context length of 3 or 4 words, while LLMs have a context length of many thousand words. Your required amounts of training data would go to "number of atoms in the universe" territory if you wanted to create a Markov chain with comparable context length.
Secondly, as current LLMs are based on the mathematical abstraction of neural networks, the relationship between training data and the eventual model weights/parameters isn't even fully deterministic: The weights are set to initial values based on some process that is independent of the training data - e.g. they are set to random values - and then incrementally adjusted so that the model can increasingly replicate the training data. This means that the "meaning" of individual weights and their relationship to the training data remains very unclear, and there is plenty of space in the model where higher-level "semantic" representations might evolve.
None of that is proof that LLMs have "intelligence", but I think it does show that the question can't be simply dismissed by saying that LLMs are statistical models.