The issue is that humans don't talk like this. I don't ask someone how many r's there are in strawberry by spelling out strawberry, I just say the word.
The issue is that humans don't talk like this. I don't ask someone how many r's there are in strawberry by spelling out strawberry, I just say the word.
Especially native speakers are prone to the mistake as they grew up learning english as illiterate children, from sounds only, compared to how most people learning english as second language do it, together with the textual representation.
Psychologists use this trick as well to figure out internal representations, for example the rorschach test.
And probably, if you asked random people in the street how many p's there is in "Philippines", you'd also get lots of wrong answers. It's tricky due to the double p and the initial p being part of an f sound. The demonym uses "F" as the first letter, and in many languages, say Spanish, also the country name uses an F.
No, I would actually be pretty confident you don’t ask people that question… at all. When is the last time you asked a human that question?
I can’t remember ever having anyone in real life ask me how many r’s are in strawberry. A lot of humans would probably refuse to answer such an off-the-wall and useless question, thus “failing” the test entirely.
A useless benchmark is useless.
In real life, people overwhelmingly do not need LLMs to count occurrences of a certain letter in a word.
If full artificial intelligence, as we're being promised, falls short in this simple way.
Problems can exist as instances of a class of problems. If you can't solve a problem, it's useful to know if it's a one off, or if it belongs to a larger class of problems, and which class it belongs to. In this case, the strawberry problem belongs to the much larger class of tokenization problems - if you think you've solved the tokenization problem class, you can test a model on the strawberry problem, with a few other examples from the class at large, and be confident that you've solved the class generally.
It's not about embodied human constraints or how humans do things; it's about what AI can and can't do. Right now, because of tokenization, things like understanding the number of Es in strawberry are outside the implicit model of the word in the LLM, with downstream effects on tasks it can complete. This affects moderation, parsing, generating prose, and all sorts of unexpected tasks. Having a workaround like forcing the model to insert spaces and operate on explicitly delimited text is useful when affected tasks appear.
We are also not exactly looking letter by letter at everything we read.
However, this implies you need to know the word to begin with.
I can write "asdf" and you might be oblivious to what I mean. I can mention "adsf" to a JavaScript developer and he will immediately think of the tool versioning tool. Because context and familiarity is important.
"vox"
"il"
"ha"
"pen"
"tosh"
is how I read it.
A lot of schools teach kids to read with a syllabic method... so... super close to the tokenization concept.
https://chatgpt.com/share/4298efbf-1c29-474a-b333-c6cc1a3ce3...
A token is roughly four letters [1], so, among other probable regressions, this would significantly reduce the effective context window.
[1] https://help.openai.com/en/articles/4936856-what-are-tokens-...
User: Write “strawberry” one letter at a time, with a space between each letter. Then count how many r’s are in strawberry.
gpt-3.5-turbo: ASSISTANT s t r a w b e r r y
There are 2 r's in strawberry.
After some experimenting, it seems like the actual problem is that many LLMs can’t count.
User: How many r’s are in the following sequence of letters:
S/T/R/A/W/B/E/R/R/Y
gpt-4o-mini: In the sequence S/T/R/A/W/B/E/R/R/Y, there are 2 occurrences of the letter "R."
Oddly, if I change a bunch of the non-R letters, I seem to start getting the right answer.
You don't ask a human being how many r's there are in strawberry at all. The only reason you or anyone else asks that question is because it's an interesting quirk of how LLMs work that they struggle to answer it in that format. It's like an alien repeatedly showing humans an optical illusion that relies on the existence of our (literal) blind spot and using it as evidence of our supposed lack of intelligence.
Yes, it's an issue. We want the convenience of sending human-legible commands to LLMs and getting back human-readable responses. That's the entire value proposition lol.