LLMs don't actually "see" individual input characters, they see tokens, which are subwords. As far as they can "see", tokens are indivisible, since the LLM doesn't get access to individual characters at all. So it's impossible for them to count letters natively. Of course, they could still get the question right in an indirect way, e.g. if a human at some point wrote "strawberry has three r's" and this text ends up in the LLM's training set, it could just use that information to answer the question just like they would use "Paris is the capital of France" or whatever other facts they have access to. But they can't actually count the letters, so they are obviously going to fail often. This says nothing about their intelligence or reasoning capability, just like you wouldn't judge a blind person's intelligence for not being able to tell if an image is red or blue.
On the other hand, writing code to count appearances of a letter doesn't run into the same limitation. It can do it just fine. Just like a blind programmer could code a program to tell if an image is red or blue.
I would judge a blind person's intelligence if they couldn't remember the last sentence they spoke when specifically asked. Or if they couldn't identify how many people were speaking in a simple audio dialogue.
This absolutely says something about their intelligence or reasoning capability. You have this comment:
> LLMs don't actually "see" individual input characters, they see tokens, which are subwords.
This alone is an indictment of their "reasoning" capability. People are saying these models understand theoretical physics but can't do what a 5 year old can do in the medium of text. It means that these are very much memorization/interpolation devices. Anything approximating reasoning is stepping through interpolation of tokens (and not even symbols) in the text. It means they're a runaway energy minimization algorithm chained to a set of tokens in their attention window, without the ability to reflect upon how any of those words relate to each other outside of syntax and ordering.
I'm not sure why it says anything about their reasoning capability. Some people are blind and can't see anything. Some people are short-sighted and can't see objects which are too far away. Some people have dyslexia. Does it say anything about their reasoning capability?
LLMs "perceive" the world through tokens just like blind people perceive the world through touch or sound. Blind people can't discuss color just like LLMs can't count letters. I'm not saying LLM's can actually reason, but I think a different way to perceive the world says nothing about your reasoning capability.
Did humans acquire reasoning capabilities only after the invention of the alphabet? A language isn't even required to have an alphabet, see Chinese. The question "how many letters in word X" doesn't make any sense in Chinese. There are character-level LLMs which can see every individual letter, but they're apparently less efficient to train.
The fact they can't operate on full symbols reliably but require sub-symbols via tokens is concrete proof of that. They may add heuristics or build more CoT sub-chains to get around some of these trickier issues later, but this is the state of affairs right now.
All efforts so far require exponential increases in training size to receive logarithmic increases (at best) in accuracy. And now with o1, it requires exponential compute at inference to scale with that sub-logarithmic accuracy.
People have a short memory these days, but around GPT-3, the majority of people on HN and tech "luminary" founders were saying that these would actually have exponential output and diverge. They were wrong. These models are quickly converging to a training set because they are and always were a curve fit. And even there, they are notoriously unreliable for use cases without a human in the loop, because of the intrinsic amount of information entropy that can be packed into the size of these models. But there is nothing mysterious about them.
Perhaps it's an activation issue (i.e. broken after all) and it just needs an occasional change of basis.
Is it? These stupid word generators are marketed as AI, I don't think it's "shocking" that people think something "intelligent" could perform a trivial counting task. My 6 year old nephew could solve it very easily.
And if you forgo the counting and just ask it to list the letters it is almost always correct, even though, once again, it never sees the input characters.
Much has been written about how tokenization hurts tasks that the LLM providers literally market their model on (Anthropic Hiaku, Sonnet): https://aclanthology.org/2022.cai-1.2/
Even if strawberry is decomposed as "straw-berry", the required logic to calculate 1+2 seems perfectly within reach.
Also, the LLM could associate a sequence of separate characters to each token. Most LLMs can spell out words perfectly fine.
Am I missing something?
I don't understand how most LLMs can spell out words though, nor do I understand what is causing the failure to count characters in words. I was not convinced by the comment I was responding to.
You should ask why it is that any of those tasks work, rather than ask why counting letter doesn't work.
Also, LLMs screw up many of those tasks more than you'd expect. I don't trust LLMs with any kind of numeracy what-so-ever.
For example, according to https://platform.openai.com/tokenizer, "strawberry" would be tokenized by the GPT-4o tokenizer as "st" "raw" "berry" (tokens don't have to make sense because they are based on byte-pair encoding, which boils down to n-gram frequency statistics, i.e. it doesn't use morphology, syllables, semantics or anything like that).
Those tokens are then converted to integer IDs using a dictionary, say maybe "st" is token ID 4663, "raw" is 2168 and "berry" is 487 (made up numbers).
Then when you give the model the word "strawberry", it is tokenized and the input the LLM receives is [4463, 2168, 487]. Nothing else. That's the kind of input it always gets (also during training). So the model has no way to know how those IDs map to characters.
As some other comments in the thread are saying, it's actually somewhat impressive that LLMs can get character counts right at least sometimes, but this is probably just because they get the answer from the training set. If the training set contains a website where some human wrote "the word strawberry has 3 r's", the model could use that to get the question right. Just like if you ask it what is the capital of France, it will know the answer because many websites say that it's Paris. Maybe, just maybe, if the model has both "the word straw has 1 r" and "the word berry has 2 r's" and the training set, it might be able to add them up and give the right answer for "strawberry" because it notices that it's being asked about [4463, 2168, 487] and it knows about [4463, 2168] and [487]. I'm not sure, but it's at least plausible that a good LLM could do that. But there is no way it can count characters in tokens, it just doesn't see them.
If I'm missing something and you have a source for the claim that character information is present in the input after tokenization, please provide it. I have never implemented an LLM or fiddled with them at low level so I might be missing some detail, but from everything I have read, I'm pretty sure it doesn't work that way.
> But there is no way it can count characters in tokens, it just doesn't see them.
If that is the case, then how can most LLMs (tested with ChatGPT and Llama 3) spell out words correctly?
One can define "reasoning" in the context of AI as the ability to perform logic operations in a loop with decisions to arrive at an answer. LLMs can't really do this.
Uh I'm sorry but I think it's not as easy as it seems. A pixel? Sure it's easy just compare whether the blue is bigger than red value. For image, I don't think it's as easy.