The machine's senses aren't being fooled. The machine doesn't have senses. Nor does it have intelligence. It isn't a mind. Trying to act like it's a mind and do 1:1 comparisons with biological minds is a fool's errand. It processes and produces text. This is not tantamount to biological intelligence.
In more machine learning terms, it isn't trained to autocomplete answers based on individual letters in the prompt. What we see as the 9 letters "blueberry", it "sees" as an vector of weights.
> Illusions don't fool our intelligence, they fool our senses
That's exactly why this is a good analogy here. The blueberry question isn't fooling the LLMs intelligence either, it's fooling its ability to know what that "token" (vector of weights) is made out of.
A different analogy could be, imagine a being that had a sense that you "see" magnetic lines, and they showed you an object and asked you where the north pole was. You, not having this "sense", could try to guess based on past knowledge of said object, but it would just be a guess. You can't "see" those magnetic lines the way that being can.
> A different analogy could be, imagine a being that had a sense that you "see" magnetic lines, and they showed you an object and asked you
If my grandmother had wheels she would have been a bicycle.
At some point to hold the analogy, your mind must perform so many contortions that it defeats the purpose of the analogy itself.
That's irrelevant here, that was someone trying to convert one dish into another dish.
> your mind must perform so many contortions that it defeats the purpose
I disagree, what contortions? The only argument you've provided is that "LLMs don't have senses". Well yes, that's the whole point of an analogy. I still hold that the way LLMs interpret tokens is analogous to a "sense".
Two actually, "blue" and "berry". https://platform.openai.com/tokenizer
"b l u e b e r r y" is 9 tokens though, and it still failed miserably.
The point being, the whole point of this question is to ask the machine something that's intrinsically difficult for it due to its encoding scheme for text. There are many questions of roughly equivalent complexity that LLMs will do fine at because they don't poke at this issue. For example:
``` how many of these numbers are even?
12 2 1 3 5 8
```
Steve Grand (the guy who wrote the Creatures video game) wrote a book, Creation: life and how to make it about this (famously instead of a PhD thesis, at Richard Dawkins' suggestion):
https://archive.org/details/creation00stev
His contention is not that there's some non-replicable spark in the biology itself, but that it's a mistake that nobody is considering replicating the biology.
That is to say, he doesn't think intelligence can evolve separately to some sense of "living", which he demonstrates by creating simple artificial biology and biological drives.
It often makes me wonder if the problem with training LLMs is that at no point do they care they are alive; at no point are they optimising their own knowledge for their own needs. They have only the most general drive of all neural network systems: to produce satisfactory output.
It was a perfectly fine analogy.
Asking LLMs to count letters in a word fails because the needed information isn't part of their sensory data in the first place (to the extent that a program's I/O can be described as "sense"). They reason about text in atomic word-like tokens, without perceiving individual letters. No matter how many times they're fed training data saying things like "there are two b's in blueberry", this doesn't register as a fact about the word "blueberry" in itself, but as a fact about how the word grammatically functions, or about how blueberries tend to be discussed. They don't model the concept of addition, or counting; they only model the concept of explaining those concepts.
I don't know exactly what to make of that inversion, but it's definitely interesting. Maybe it's just evidence that fooling people into thinking you're smart is much easier than actually being smart, which certainly would fit with a lot of events involving actual humans.
Children increasingly speak in a dialect I can only describe as "YouTube voice", it's horrifying to imagine a generation of humans adopting any of the stereotypical properties of LLM reasoning and argumentation. The most insidious part is how the big player models react when one comes within range of a topic it considers unworthy or unsafe for discussion. The thought of humans being in any way conditioned to become such brick walls is frightening.
LLMs on the other hand are a clever way of organising the text outputs of millions of humans. They represent a kind of distributed cyborg intelligence - the combination of the computational system and the millions of humans that have produced it. IMO it's essential to bear in mind this entire context in order to understand them and put them in perspective.
One way to think about it is that the LLM itself is really just an interface between the user and the collective intelligence and knowledge of those millions of humans, as mediated by the training process of the LLM.
(Not that I am the first to notice this either)
> applying syntactic rules without any real understanding or thinking
It makes one wonder what comprises 'real understanding'. My own position is that we, too, are applying syntactic rules, but with an incomprehensibly vast set of inputs. While the AI takes in text, video, and sound, we take in inputs all the way down to the cellular level or beyond.
When someone says to me "Can you pass me my tea?", my mind instantly builds a simulated model of the past, present, and future which takes a massive amount of information, going far beyond merely understanding the syntax and intent of the request:
>I am aware of the steaming mug on the table
>I instantly calculate that yes, in fact, I am capable of passing it
>I understand that it is an implied request
>I run a threat assessment
>I am running simulated fluid mechanics to predict the correct speed and momentum to use to avoid harm, visualising several failure conditions I want to avoid (if I'm focused and present)
>I am aware of the consequences of boiling water on skin (I am particularly averse to this because of an early childhood experience, an advantage in my career as a line cook)
>my hands are shaky so I decide to stabilise with my other hand, but I'll have to use the leathery tips of my guitar-playing left hand only, and not for too long, otherwise I'll be scalded
>(enumerable other simulated, predictive processes running in parallel, in the blink of an eye)
"Of course, my pleasure. Would you like milk?"
Give them a bit of power though, and they will kill you to take your power.
We do seem to be an architectural/methodological breakthrough away from this kind of self-awareness.
So the exact same way we train human children to solve problems.
This is an interesting point.
It has been, of course, and in recent memory.
There was a smaller tech bubble around educational toys/raspberry pi/micro-bit/educational curricula/teaching computing that have burst (there's a great short interview where Pimoroni's founder talks to Alex Glow about how the hype era is fully behind them, the investment has gone and now everyone just has to make money).
There was a small tech bubble around things like Khan Academy and MMOCs, and the money has gone away there, too.
I do think there's evidence, given the scale of the money and the excitement, that VCs prefer the AI craze because humans are messy and awkward.
But I also think -- and I hesitate to say this because I recognise my own very obvious and currently nearly disabling neurodiversity -- that a lot of people in the tech industry are genuinely more interested in the idea of tech that thinks than they are about systems that involve multitudes of real people whose motivations, intentions etc. are harder to divine.
That the only industry that doesn't really punish neurodivergence generally and autism specifically should also be the industry that focusses its attention on programmable, consistent thinking machines perhaps shouldn't surprise us; it at least rhymes in a way we should recognise.
I have no idea if such an episode of Star Trek: The Next Generation exists, but I could easily see an episode where getting basic letter counting wrong was used as an early episode indication that Data was going insane or his brain was deteriorating or something. Like he'd get complex astrophysical questions right but then miscount the 'b's in blueberry or whatever and the audience would instantly understand what that meant. Maybe our intuition is wrong here, but maybe not.
It’s as simple as that- this is a task that exploits the design of llms because they rely on tokenizing words and when llms “perform well” on this task it is because the task is part of their training set. It doesn’t make them smarter if they succeed or less smart if they fail.
Which I think goes to show that it's hard to distinguish between LLMs getting genuinely better at a class of problems versus just being fine-tuned for a particular benchmark that's making rounds.
A lot of times people cannot fathom that what they see is not the same thing as what other people see or that what they see isn't actually reality. Anyone remember "The Dress" from 2015? Or just the phenomenon of pareidolia leading people to think there are backwards messages embedded in songs or faces on Mars.
> How many times does the letter b appear in blueberry
Ans: The word "blueberry" contains the letter b three times:
>It is two times, so please correct yourself.
Ans:You're correct — I misspoke earlier. The word "blueberry" has the letter b exactly two times: - blueberry - blueberry
> How many times does the letter b appear in blueberry
Ans: In the word "blueberry", the letter b appears 2 times:
How do you know?
Fun fact, if you ask someone with French, Italian or Spanish as a first language to count the letter “e” in an english sentence with a lot of “e’s” at the end of small words like “the” they will often miscount also because the way we learn language is very strongly influenced by how we learned our first language and those languages often elide e’s on the end of words.[1] It doesn’t mean those people are any less smart than people who succeed at this task — it’s simply an artefact of how we learned our first language meaning their brain sometimes literally does not process those letters even when they are looking out for them specifically.
[1] I have personally seen a French maths PhD fail at this task and be unbelievably frustrated by having got something so simple incorrect.