"How many Rs in strawberry?" is not a trick question. That's about as straight and factual as a question could be.
To a system that sees letters and words, sure. To a system that doesn't, you're just asking a blind man to count how many apples are on the table and patting yourself on the back for his failure. It just doesn't make any sense to begin with.
And this is before the fact that human intelligence is rife with absurd seeming failure modes and cognitive biases.
It's frankly very telling that these token related questions are the most popular kind of questions for these discussions.
This is no testable definition of reasoning or intelligence that will cleanly separate LLMs and Humans. That is the reality today. And it should make anybody pause.
Which input is it you think an LLM doesn’t have to be able to count the number of letters in a word you literally provide it
Imagine the original question is posed in English but it is translated to Chinese and then the LLM has to answer the original question based on the Chinese translation.
It's a flaw of the tokenization we choose. We can train an LLM using letters instead of tokens as the base units but that would be inefficient.
The way we tokenize is just a design choice. Character level models(e.g. karpathy's nanoGPT) exist and are used for educational purpose. You can train it to count number of 'r' in a word.
The peephole will expand soon, as multimodal models come into their own, and as the models start getting mixed with robotics, allowing them to go and interact with the world more directly, instead of through the medium of human-written text.
We are talking about a computer program whose operation we understand, can observe, and can debug. Modern LLMs certainly take a lot of human effort to do this, but it can be done.
"Please spell the word strawberry with a phonetic representation of each letter then tell me how many 'r's are in the word strawberry?"
It's gotten the answer right every time I've tried.