Can modern LLMs count the number of b's in "blueberry"?
minimaxir.com
minimaxir.com
Ask Jeeves from 1997 could answer this question, so tell me why we need to devote a nation-state amount of compute power to feed an “AI” that confidently gets kindergarten level questions dead ass wrong?
I have the same kind of question when I watch the AI summary on Google output tokens one-by-one to give me less useful information that is right there on the first search result from Wikipedia (fully sourced, too)
The author tested 12 models, and only one was consistently wrong. More than half were correct 100% of the time.
A better conclusion would be that there’s something in particular wrong with GPT-5 Chat, all the other GPT 5 variants are OK. I wonder what’s different?
This is why Anthropic naming system of haiku sonnet and opus to represent size is really nice. It prevents this confusion.
In contrast to GPT-5, GPT-5 mini and GPT-5 nano?
4.large, 4.medium, 4.fast, 4. reasoning etc. or something similar would probably be better.
Anthropic model names might not immediately conjure up their size and performance, but the name is at least internally consistent. Once you know what Anthropic call “medium”, you know what it is for all model releases.
Whereas OpenAIs naming convention, if you can even call it a “convention”, feels absolutely random to even those in the industry.
I do like your proposed naming convention though. It doesn’t sound “cool” so I can’t see any product managers approving it within the AI tech firms. But it’s definitely the best naming convention for models I’ve seen suggested for a while.
If anything it's a lot less confusing that the awful naming convention from OpenAI up until 5.
Sure, an opus is supposed to be large, but a sonnet is not restricted in size but rather a style of poem. So sonnet and opus mean nothing when compared to each other.
It might not be perfect, but it’s still a hell of a lot better.
Also, some ChatGPT models include “gpt” in the name. Others do not.
I cannot guess what model string I need to pass. Whereas with Anthropic I can. And if I have to look it up each time on OpenAIs website, then it’s clearly garbage.
Also the “arcane barely used” part of your post is entirely subjective. I get you want to make the point that Anthropic naming is poor to support your point about OpenAI, but you’re over exaggerating your point there.
I would frequently spend time going back to OpenAIs site to remind me of their different models. There’s no consistency there whatsoever. But with Anthropic is was easy.
If I have to spend 5 minutes picking a model then I might as well do the task myself. So Claude became a natural solution for me.
For a certain class of customer maybe that is true.
But the reality is that the fact that this occurs is very encouraging -- they are not micro-optimizing to solve cosmetic problems that serve no functional purpose. They are instead letting these phenomena serve as external benchmarks of a sort to evaluate how well the LLM can work on tasks that are outside of its training data, and outside of what one would expect the capabilities to be.
It doesn't seem valuable for the model to memorize the graphemes in each of its tokens.
The reason LLMs struggle with this is because they literally aren't thinking in English. Their input is tokenized before it comes to them. It's like asking a Chinese speaker "How many Rs are there in the word 草莓".
Since tokens are atomic, which I didn't realize earlier, then maybe it's still intelligent if it can realize it can extract the result by writing len([b for b in word if b == my_letter]) and decide on its own to return that value.
spoiler: Yes. This has already become standard for production use cases where the LLM is an external-facing interface; you use an LLM to translate the user's human-language request to a machine-ready, well-defined schema (i.e. a protobuf RPC), do the bulk of the actual work with actual, deterministic code, then (optionally) use an LLM to generate a text result to display to the user. The LLM only acts as a user interface layer.
Turned out it just knew that particular string because it was on the internet. I don’t recall exactly but it was probably “hello world” or something like that.
Anyway to your specific question, many of the models call Python tools behind the scenes now.
GPT-5: "How many times does the letter b appear in blueberry?" - https://news.ycombinator.com/item?id=44832908 - Aug 2025 (321 comments)
(My one-off test of the default ChatGPT model, whatever that is, got 'How many b's are there in "Billy Bob beat the record for bounding?"' correct first try, with correct reasoning given.)
ollama run qwen3:4b-instruct-2507-fp16 >>> how many b’s are there in blueberry? Let's break down the word *"blueberry"* step by step to count how many *b's* are in it.
Step 1: Write out the word clearly: *b l u e b e r r y*
Step 2: Go through each letter one by one:
- b → yes, a *b* - l → no - u → no - e → no - b → yes, another *b* - e → no - r → no - r → no - y → no
Step 3: Count the *b's*:
- First *b* at the beginning - Second *b* after "blue"
So, there are *two b's*.
Final answer: *2* b's in "blueberry".
>>> ollama run hf.co/ibm-granite/granite-3.3-2b-instruct-GGUF:F16 >>> how many b’s are there in blueberry? The word "blueberry" contains two 'b's. (fastest lol, granite models are pretty underated)
r1-distill output was similar to qwen instruct one but it double checked it's thinking part
The AI thought and concluded that he had lost his job first, until I pointed out that it was not the first thing he had lost - which was his umbilical cord, a far better answer, in the AI's opinion.
Which raises many aspects - Can an AI disagree with you? Will AI develop solid out-of-the-box thinking as well as in-the-box thinking, will it grasp applying both for a thru the box thinking and solutions...
After all, we have yet to perfect the teaching of children, so the training of AI, has a long way to go and will get down to quality over quantity, just deciding what is quality and what is not. After all - Garbage in, Garbage out, is probably more important today than it ever was in the history of technology.
GPT-5: "How many times does the letter b appear in blueberry?"
This is a great example. The LLM doesn't know something but it makes up something in it's place. Just because it made up something doesn't mean it's incapable of reasoning.
The thing with LLMs is that they can reason. There's evidence for that. But they can also be creative. And the line between reasoning and creativity at a low level is a bit of a blur as reasoning is a form of inference but so is creativity. So when an LLM reasons or gets creative or hallucinates it's ultimately doing the same type of thing: inference.
For us, we have mechanisms in our brain that allow us to tell the difference most of the time. The LLM does not. That's the fundamental line. And I feel because of this we are literally really close to AGI. A lot of people argue the opposite. They think reasoning and is core to intelligence and a separate concept from creativity and that all LLMs lack reasoning. I disagree.
In fact humans ourselves have trouble separating hallucination from reasoning. Look at religion. Religion permeates our culture but it's basically all hallucinations that we ultimately mistake for reasoning. Right? Ask any christian or muslim, the religions make rational sense to them! They can't tell the difference.
So the key is to give the LLM the ability to know the difference.
Is there some way to build into the transformer, some way to quantify whether something is fact or fiction? Like let's say the answer to a prompt created an inferenced datapoint that's very far far away from a cluster of data. From that we can derive some metric that quantifies how likely the response is based on evidence?
Right? The whole thing is on a big mathematical multidimensional durve. If the inferenced point on the curve is right next to existing data then it must be more likely to be true. If it's far away in some nether region of the curve then it's more likely to be false.
If the LLM can be more self aware and we can build this quantitative metric into the network then use reinforcement learning to kind of have the network be less sure about an answer if it's far away from a cluster of training data points we can likely very much improve the hallucination problem.
Of course I'm sure this is a blunt instrument as even false inferences data can be very close to existing training data. But at least this gives the LLM some level of self awareness of how reliable it's own answer was.
Dev: “What about Bs in blueberry?”
PM: “you’ll need to open a new jira ticket”
A non-English speaking adult wouldn't be able to answer the question either, after the question was translated to their language. Maybe you wouldn't want a non-English speaker helping you to write an acrostic in English. Luckily nobody is marketing LLMs as "great for designing word puzzles" though.
That is literally useless knowledge.
The point is that LLMs don't speak / think in English. Asking them about spelling is like asking a Chinese speaker, during a text chat with translation, about English spelling. We can give the Chinese speaker access to an app to translate back to English so they can answer these questions. But they (the LLM) don't currently have access to that.
They don't think or speak anything at all. They use a statistical model to predict the next most likely token to display in response to a prompt.
s/speak \/ think/operate/
If that makes you happy.