See
See
Tokenization makes the problem difficult, but not solving it is still a reasoning/intelligence issue
> How many "s"es are in the word "Mississippi"?
The "thinking portion" is:
> Count letters: M i s s i s s i p p i -> s appears 4 times? Actually Mississippi has s's: positions 3,4,6,7 = 4.
The answer is:
> The word “Mississippi” contains four letter “s” s.
They can indeed do some simple pattern matching on the query, separate the letters out into separate tokens, and count them without having to do something like run code in a sandbox and ask it the answer.
The issue here is just that this workaround/strategy is only trained into the "thinking" models, afaict.
And now that fact is going to be in the data for the next round of training. We'll need to need to try some other words on the next model.
(Sometimes the trace is noisier, especially in quants other than the original.)
This task is pretty simple and I think can be solved easily with the same kind of statistical pattern matching these models use to write other text.
(Yes, I'm sure an agentic + "reasoning" model can already deduce the strategy of writing and executing a .count() call in Python or whatever. That's missing the point.)
https://claude.ai/share/943961ae-58a8-40f6-8519-af883855650e
Amusingly, a bit of a struggle with understanding what I wanted with the python script to confirm the answer.
I really don't get why people think this is some huge un-fixable blindspot...
Nobody who could give answers as good as ChatGPT often does would struggle so much with this task. The fact that an LLM works differently from a whole-ass human brain isn't actually surprising when we consider it intellectually, but that habit of always intuiting a mind behind language whenever we see language is subconscious and and reflexive. Examples of LLM failures which challenge that intuition naturally stand out.
For GPT 5, it would seem this depends on which model your prompt was routed to.
And GPT 5 Thinking gets it right.
This matters because it poses a big problem for the (quite large) category of things where people expect LLMs to be useful when they get just a bit better. Why, for example, should I assume that modern LLMs will ever be able to write reliably secure code? Isn’t it plausible that the difference between secure and almost secure runs into some similar problem?
Have you got any proof they're even trying? It's unlikely that's something their real customers are paying for.
It's not in their interest to write off the scheme as provably unworkable at scale, so they keep working on the edge cases until their options vest.
If you're fine appealing to less concrete ideas, transformers are arbitrary function approximators, tokenization doesn't change that, and there are proofs of those facts.
For any finite-length function (like counting letters in a bounded domain), it's just a matter of having a big enough network and figuring out how to train it correctly. They just haven't bothered.
Or they don't see the benefit. I'm sure they could train the representation of every token and make spelling perfect. But if you have real users spending money on useful tasks already - how much money would you spend on training answers to meme questions that nobody will pay for. They did it once for the fun headline already and apparently it's not worth repeating.
You seem to suppose that they actually perform addition internally, rather than simply having a model of the concept that humans sometimes do addition and use it to compute results. Why?
> For any finite-length function (like counting letters in a bounded domain), it's just a matter of having a big enough network and figuring out how to train it correctly. They just haven't bothered.
The problem is that the question space grows exponentially in the length of input. If you want a non-coincidentally-correct answer to "how many t's in 'correct horse battery staple'?" then you need to actually add up the per-token counts.
Nothing of the sort. They're _capable_ of doing so. For something as simple as addition you can even hand-craft weights which exactly solve it.
> The problem is that the question space grows exponentially in the length of input. If you want a non-coincidentally-correct answer to "how many t's in 'correct horse battery staple'?" then you need to actually add up the per-token counts.
Yes? The architecture is capable of both mapping tokens to character counts and of addition with a fraction of their current parameter counts. It's not all that hard.
It's really disingenuous for the industry to call warming tokens for output, "reasoning," as if some autocomplete before more autocomplete is all we needed to solve the issue of consciousness.
Edit: Letter frequency apparently has just become another scripted output, like doing arithmetic. LLMs don't have the ability to do this sort of work inherently, so they're trained to offload the task.
Edit: This comment appears to be wildly upvoted and downvoted. If you have anything to add besides reactionary voting, please contribute to the discussion.
Of course then you ask her to write it and of course things get fixed. But strange.
That is to say, you can obtain the same process by talking to "non-reasoning" models.
There will be a series of analytical articles in the mainstream press, the tech industry will write it off as a known problem with tokenisation that they can't fix because nobody really writes code anymore.
The LLM megacorp will just add a disclaimer: the software should not be used in legal actions concerning fruit companies and they disclaim all losses.
Mechanistic research at the leading labs has shown that LLMs actually do math in token form up to certain scale of difficulty.
> This is a real-time, unedited research walkthrough investigating how GPT-J (a 6 billion parameter LLM) can do addition.
There's prior art for formal logic and knowledge representation systems dating back several decades, but transformers don't use those designs. A transformer is more like a search algorithm by comparison, not a logic one.
That's one issue, but the other is that reasoning comes from logic, and the act of reasoning is considered a qualifier of consciousness. But various definitions of consciousness require awareness, which large language models are not capable of.
Their window of awareness, if you can call it that, begins and ends during processing tokens, and outputting them. As if a conscious thing could be conscious for moments, then dormant again.
That is to say, conscious reasoning comes from awareness. But in tech, severing the humanities here would allow one to suggest that one, or a thing, can reason without consciousness.
The hard truth is we have no idea. None. We got ideas and conjectures, maybe's and probably's, overconfident researchers writing books while hand waving away obvious holes, and endless self introspective monologues.
Don't waste your time here if you know what reasoning and consciousness are, go get your nobel prize.
Wrong, it's an artifact of tokenizing. The model doesn't have access to the individual letters, only to the tokens. Reasoning models can usually do this task well - they can spell out the word in the reasoning buffer - the fact that GPT5 fails here is likely a result of it incorrectly answering the question with a non-reasoning version of the model.
> There's no real reasoning.
This seems like a meaningless statement unless you give a clear definition of "real" reasoning as opposed to other kinds of reasoning that are only apparant.
> It seems that reasoning is just a feedback loop on top of existing autocompletion.
The word "just" is doing a lot of work here - what exactly is your criticism here? The bitter lesson of the past years is that relatively simple architectures that scale with compute work surprisingly well.
> It's really disingenuous for the industry to call warming tokens for output, "reasoning," as if some autocomplete before more autocomplete is all we needed to solve the issue of consciousness.
Reasoning and consciousness are seperate concepts. If I showed the output of an LLM 'reasoning' (you can call it something else if you like) to somebody 10 years ago they would agree without any doubt that reasoning was taking place there. You are free to provide a definition of reasoning which an LLM does not meet of course - but it is not enough to just say it is so. Using the word autocomplete is rather meaningless name-calling.
> Edit: Letter frequency apparently has just become another scripted output, like doing arithmetic. LLMs don't have the ability to do this sort of work inherently, so they're trained to offload the task.
Not sure why this is bad. The implicit assumption seems to be that an LLM is only valueable if it literally does everything perfectly?
> Edit: This comment appears to be wildly upvoted and downvoted. If you have anything to add besides reactionary voting, please contribute to the discussion.
Probably because of the wild assertions, charged language, and rather superficial descriptions of actual mechanics.
> Reasoning and consciousness are seperate(sic) concepts
No, they're not. But, in tech, we seem to have a culture of severing the humanities for utilitarian purposes, but no, classical reasoning uses consciousness and awareness as elements of processing.
It's only meaningless if you don't know what the philosophical or epistemological definitions of reasoning are. Which is to say, you don't know what reasoning is. So you'd think it was a meaningless statement.
Do computers think, or do they compute?
Is that a meaningless question to you? I'm sure given your position it's irrelevant and meaningless, surely.
And this sort of thinking is why we have people claiming software can think and reason.
> "classical reasoning uses consciousness and awareness as elements of processing"
They are not the _same_ concept then.
> It's only meaningless if you don't know what the philosophical or epistemological definitions of reasoning are. Which is to say, you don't know what reasoning is. So you'd think it was a meaningless statement.
The problem is the only information we have is internal. So we may claim those things exist in us. But we have no way to establish if they are happening in another person, let alone in a computer.
> Do computers think, or do they compute?
Do humans think? How do you tell?
> No, they're not. But, in tech, we seem to have a culture of severing the humanities for utilitarian purposes [...] It's only meaningless if you don't know what the philosophical or epistemological definitions of reasoning are.
As far as I'm aware, in philosophy they'd generally be considered different concepts with no consensus on whether or not one requires the other. I don't think it can be appealed to as if it's a settled matter.
Personally I think people put "learning", "reasoning", "memory", etc. on a bit too much of a pedestal. I'm fine with saying, for instance, that if something changes to refine its future behavior in response to its experiences (touch hot stove, get hurt, avoid in future) beyond the immediate/direct effect (withdrawing hand) then it can "learn" - even for small microorganisms.
I like to say that if regular LLM "chats" are actually movie scripts being incrementally built and selectively acted-out, then "reasoning" models are a stereotypical film noir twist, where the protagonist-detective narrates hidden things to himself.
There's no obvious connection between reasoning and consciousness. It seems perfectly possible to have a model that can reason without being conscious.
Also, dismissing what these models do as "autocomplete" is extremely disingenuous. At best it implies you're completely unfamiliar with the state of the art, at worst it implies an dishonest agenda.
In terms of functional ability to reason, these models can beat a majority of humans in many scenarios.
A locally trained text-based foundation model is indistinguishable from autocompletion, and outputs very erratic text, and the further you train it's ability to diminish irrelevant tokens, or guide it to produce specifically formatted output, you've just moved its ability to curve fit specific requirements.
So it may be disingenuous to you, but it does behave very much like a curve fitting search algorithm.
Unless you can show us that humans can calculate functions outside the Turing computable, it is logical to conclude that computers can be made to think due to Turing equivalence and the Church Turing thesis.
Given we have zero evidence to suggest we can exceed the Turing computable, to suggest we can is an extraordinary claim that requires extraordinary evidence.
A single example of a function that exceeds the Turing computable that humans can compute, will do.
Until you come up with that example, I'll asume computer can be made to think.
What matters here is a functional definition of reasoning: something that can be measured. A computer can reason if it can pass the same tests that humans can pass of reasoning ability. LLMs blew past that milestone quite a while back.
If you believe that "thinking" and "reasoning" have some sort of mystical aspect that's not captured by such tests, it's up to you to define that. But you'll quickly run into the limits of such claims, because if you want to attribute some non-functional properties to reasoning or thinking, that can't be measured, then you also can't prove that they exist. You quickly get into an intractable area of philosophy, which isn't really relevant to the question of what AI models can actually do, which is what matters.
> it does behave very much like a curve fitting search algorithm.
This is just silly. I can have an hours-long coding session with an LLM in which it exhibits a strong functional understanding of the codebase its working on, a strong grasp of the programming language and tools its working with, and writes hundreds or thousands of lines of working code.
Please plot the curve that it's fitting in a case like this.
If you really want to stick to this claim, then you also have to acknowledge that what humans do is also "behave very much like a curve fitting search algorithm." If you disagree, please explain the functional difference.
How many 538 do you see in 423, 4144, 9890?