These discussions often get derailed into debates about what "thinking" means. If we define thinking as the capacity to produce and evaluate arguments, as the cognitive scientists Sperber and Mercier do, then we can see LLMs are certainly producing arguments, but they're weak at the evaluation.
In some cases, arguments can be formalised, and then evaluating them is a solved problem, as in the examples of using the Lean proofchecker in combination with LLMs to write mathematical proofs.
That suggests a way forward will come from formalising natural language arguments. So LLMs by themselves might be a dead end but in combination with formalisation they could be very powerful. That might not be "thinking" in the sense of the full suite of human abilities that we group with that word but it seems an important component of it.
If by this you mean "reliably convert expressions made in human natural language to unambiguous, formally parseable expressions that a machine can evaluate the same way every time"... isn't that essentially an unreachable holy grail? I mean, everyone from Plato to Russell and Wittgenstein struggled with the meaning of human statements. And the best solution we have today is to ask the human to restrict the set of statement primitives and combinations that they can use to a small subset of words like "const", "let foo = bar", and so on.
In the test setup, the AI added a single database row, ran the query and then asserted the single added row was returned. Clearly this doesn't show that the query works as intended. Is this what people are referring to when they say AI writes their tests?
I don't know what to call this kind of thinking. Any intelligent, reasoning human would immediately see that it's not even close to enough. You barely even need a coding background to see the issues. AI just doesn't have it, and it hasn't improved in this area for years
This kind of thing happens over and over again. I look at the stuff it outputs and it's clear to me that no reasoning thing would act this way
The tooling in the Code tools is key to useable LLM coding. Those tools prompt the models to “reason” whether they’ve caught edge cases or met the logic. Without that external support they’re just fancy autocompletes.
In some ways it’s no different than working with some interns. You have to prompt them to “did you consider if your code matched all of the requirements?”.
LLMs are different in that they’re sorta lobotomized. They won’t learn from tutoring “did you consider” which needs to essentially be encoded manually still.
I really hate this description, but I can't quite filly articulate why yet. It's distinctly different because interns can form new observations independently. AIs can not. They can make another guess at the next token, but if it could have predicted it the 2nd time, it must have been able to predict it the first, so it's not a new observation. The way I think through a novel problem results in drastically different paths and outputs from an LLM. They guess and check repeatedly, they don't converge on an answer. Which you've already identified
> LLMs are different in that they’re sorta lobotomized. They won’t learn from tutoring “did you consider” which needs to essentially be encoded manually still.
This isn't how you work with an intern (unless the intern is unable to learn).
That has other explanations than that it reasoned its way to the correct answers. Maybe it had very similar code in its training data
This specific example was with Codex. I didn't mention it because I didn't want it to sound like I think codex is worse than claude code
I do realize my prompt wasn't optimal to get the best out of AI here, and I improved it on the second pass, mainly to give it more explicit instruction on what to do
My point though is that I feel these situations are heavily indicative of it not having true reasoning and understanding of the goals presented to it
Why can it sometimes catch the logic cases you miss, such as in your case, and then utterly fail at something that a simple understanding of the problem and thinking it through would solve? The only explanation I have is that it's not using actual reasoning to solve the problems
yes
> Any intelligent, reasoning human would immediately see that it's not even close to enough. You barely even need a coding background to see the issues.
[nods]
> This kind of thing happens over and over again. I look at the stuff it outputs and it's clear to me that no reasoning thing would act this way
and yet there're so many people who are convinced it's fantastic. Oh, I made myself sad.
The larger observation about it being statistical inference, rather than reason... but looks to so many to be reason is quite an interesting test case for the "fuzzing" of humans. In line with why do so many engineers store passwords in clear text? Why do so many people believe AI can reason?
Hot take (and continue with the derailment), but I'd argue that analytic philosophy from the last 100 years suggests this is a dead end. The idea that belief systems could be formalized was huge in the early 20th century (movements like Logical Positivism, or Russell's principia mathematica being good examples of this).
Those approaches haven't really yielded many results, and by far the more fruitful form of analysis has been to conceptually "reframe" different problems (folks like Hillary Putnam, Wittgenstein, Quine being good examples).
We've stacked up a lot of evidence that human language is much too loose and mushy to be formalised in a meaningful way.
Lossy might also be a way of putting it, like a bad compression algorithm. Written language carries far less information than spoken and nonverbal cues.
I think modeling language usefully looks a lot more like psychoanalysis than first order logic.
"SAL-9000: Will I dream? Dr. Chandra: Of course you will. All intelligent beings dream. Nobody knows why. Perhaps you will dream of HAL... just as I often do." From 2010
That bar is so low that even a political pundit on TV can clear it.
Great. But most (?) of the business out there aren't paying for the big boy models.
I know of a F100 that got snookered into a deal with GPT 4 for 5 years, max of 40 responses per session, max of 10 sessions of memory, no backend integration.
Those folks rightly think that AI is a bad idea.
How is this sentiment still the top comment on an article about AI on HN in 2026? It's not true with today's models. They undergo vast amounts of reinforcement learning optimizing an objective that is NOT just predict the most likely next token given the training corpus. I would say even without the RL the "predict the next token" objective doesn't preclude thinking and reasoning, but that's a separate discussion. Generative sequence modeling learns to (approximately) model the process that produced the sequence. When you consider that text sequences are produced by human minds, which most would consider to be thinking and reasoning, well...
And the challenge is rethinking how we do work, connecting all the data sources for agents to run and perform work over the various sources that we perform work. That will take ages. Not to mention having the controls in place to make that the "thinking" was correct in the end.
You seem to be defining "thinking" as an interchangeable black box, and as long as something fits that slot and "gets results", it's fine.
But it's the code-writing that's the interchangeable black box, not the thinking. The actual work of software development is not writing code, it's solving problems.
With a problem-space-navigation model, I'd agree that there are different strategies that can find a path from A to B, and what we call cognition is one way (more like a collection of techniques) to find a path. I mean, you can in principle brute-force this until you get the desired result.
But that's not the only thing that thinking does. Thinking responds to changing constraints, unexpected effects, new information, and shifting requirements. Thinking observes its own outputs and its own actions. Thinking uses underlying models to reason from first principles. These strategies are domain-independent, too.
And that's not even addressing all the other work involved in reality: deciding what the product should do when the design is underspecified. Asking the client/manager/etc what they want it to do in cases X, Y and Z. Offering suggestions and proposals and explaining tradeoffs.
Now I imagine there could be some other processes we haven't conceived of that can do these things but do them differently than human brains do. But if there were we'd probably just still call it 'thinking.'
Copilot can't jump to definition in Visual Studio.
Anthropic got a lot of mileage out of teaching Claude to grep, but LLM agents are a complete dead-end for my code-base until they can use the semantic search tools that actually work on our code-base and hook into the docs for our expensive proprietary dependencies.
I have kids, and you could say the same about toddlers. Terrific mimics, they don't understand the whys.
So I think younger kids have purpose and associate meaning to a lot of things and they do try to get to a specific path toward an outcome.
Of course (depending on the age) their "reasoning" is in a different system than hours where the survival instincts are much more powerful than any custom defined outcome so most of the time that is the driving force of the meaning.
Why I talk about meaning? Because, of course, the kids cannot talk about the why, as that is very abstract. But meaning is a big part of the Why and it continues to be so in adult life it is just that the relation is reversed: we start talking about the why to get to a meaning.
I also think that kids starts to have more complex thoughts than the language very early. If you got through the "Why?" phase you might have noticed that when they ask "Why?" they could mean very different questions. But they don't know the words to describe it. Sometimes "Why?" means "Where?" sometimes means "How?" sometimes means "How long?" .... That series of questioning is, for me, a kind of proof that a lot of things are happening in kids brain much more than they can verbalise.
it’s one of the challenges when LLMs are being anthropomorphised, reasoning/logic for bots is not the same as that for humans.
But, to LLMs we don't afford the same leniency. If they flip some bits and the logic doesn't add up we're quick to point that "it's not reasoning at all".
Funny throne we've built for ourselves.
After all I can guarantee the other side (whatever it is) will say the same thing for your "logical" conclusions.
It is logic, we just don't share the same predicates or world model...
By the standard in the parent post, humans certainly do not "reason". But that is then just choosing a very high bar for "reasoning" that neither humans nor AI meets...what is the point then?
It is a bit like saying: "Humans don't reason, they just let neurons fire off one another, and think the next thought that enters their mind"
Yes, LLMs need to spew out text to move their state forward. As a human I actually sometimes need to do that too: Talk to myself in my head to make progress. And when things get just a tiny bit complicated I need to offload my brain using pen and paper.
Most arguments used to show that LLMs do not "reason" can be used to show that humans do not reason either.
To show that LLMs do not reason you have to point to something else than how it works.
If LLMs were actually able to think/reason and you acknowledge that they’ve been trained on as much data as everyone could get their hands on such that they’ve been “taught” an infinite amount more than any ten humans could learn in a lifetime, I would ask:
Why can’t they solve novel, unsolved problems?
So I will assume you mean within maths, science etc? Basically things they can't solve today.
Well 99.9% of humans cannot solve novel, unsolved problems in those fields.
LLMs cannot learn, there is just the initial weight estimation process. And that process currently does not make them good enough on novel math/theoretical physics problems.
That does not mean they do not "reason" in the same way that those 99.9% of humans still "reason".
But they definitely do not learn, the way humans do.
(Anyway, if LLMs could somehow get 1000x as large context window and get to converse with themselves for a full year, it does not seem out of the question they could come out with novel research?)
And my "reasoning" is pretty much like a long ChatGPT verbal and sometimes not-so-verbal (visual) conversation with myself.
If my mind really did abstract platonic thinking I think answers to hard problems would just instantly appear to me, without flaws. But only problems I hve solved before and can pattern match do that.
And if I have to think any new thoughts I feel that process is rather similar to how LLMs work.
It is the same for history of science really -- only thoughts that build small steps on previous thoughts and participate in a conversation actually are thought by humans.
Totally new leaps, which a "platonic thinking machines" should easily do, do not seem to happen..
Humans are, IMO, conversation machines too...
Quite the assumption you're making that the same thing does not apply to humans. Not having the technology is not the same thing as not being possible.
I see this argument made a lot. I'm not sure if the distinction really holds weight once we start to unravel though.
What's a topic you're able to think about that an LLM is not able to think about?
This is very obvious when trying to use LLMs to modify scripts in vendor-specific languages that have not been widely documented and don't have many examples available. A seasoned programmer will easily recognize common patterns like if-else blocks and loops, but LLMs will get stuck and output gibberish.
on Meta (whatsapp) chatbot:
me: what is the seahorse emoji?
bot:
me: this is not a seahorse
bot: You got me! The actual seahorse emoji is . Want me to use it in a fun sentence?
on Bing/Copilot:
me: show me the seahorse emoji
bot:
------------
Here you go: — actually, that’s a fun combo, but the real seahorse emoji is … wait, not quite! The correct one is for jellyfish and for unicorn, but the seahorse emoji itself is ? Let me clarify:
The official seahorse emoji is: ? No — correction! It’s for unicorn. The actual seahorse emoji is: ?
Let’s be precise: the seahorse emoji is ?
Oops — I misspoke earlier. The true seahorse emoji is: ?
Wait — let’s settle this cleanly: The seahorse emoji is: ?
Actually, the correct one is: ?
---------------
In some instances, Copilot answers that "there's no seahorse emoji _and that's why I keep running in circles_". This is just pasting answers, as this is a popular problem discussed on the internet. If the bot were reasoning, it would just answer "there's no seahorse emoji, please stop asking".
> Me: what is the seahorse emoji?
> Claude: There isn't a seahorse emoji in the standard Unicode emoji set. The closest you'll get is the generic fish or tropical fish , but no dedicated seahorse exists as of now.Edit: I gave up on Copilot ant fed the same instructions to ChatGPT, which had no issue.
The point here is that some models seem to know your intention while some just seem stuck on their training data.
It's an impossible burden to prove. We can't even prove that any other human has sentience or is reasoning, we just evaluate the outcomes.
One day the argument you're putting forward will be irrelevant, or good for theoretical discussion only. In practice I'm certain that machines will achieve human level output at some point.
Would you care to put some sort of time scale to "at some point?" Are we talking about months, years, decades, centuries?
Wild guess, within 30 years.