Now, of course, humans can also generate answers without reasoning too -- and in those cases, that isnt reasoning also. And in cases where people confabulate, that isnt reasoning likewise. But humans, and many classes of animals, do reason. They do reach answers via inferential entailments, not merely steered correlations.
LLMs provide imitations of arbitrary mental capacities "in the text domain", ie., the generate text as-if the LLM had those capacities. Insofar as the text generated is useful, for an engineer, that's sufficient.
As a person with scientific commitments to reality rather than its immitation, i retain the ordinary non-engineered meanings of these terms: reasoning is a deterministic inferential process over propositions; and a reasoning agent is one which has the capacity to represent propostions and their entailments, and does so when they reason. LLMs fail at all hurdles here: they have no propositonal states (ie., no rich representations), no inferential process which unites them, and so on.
You can always get abitarily close to appearing as-if, if the LLM is trained on a vast number of reasoning examples, of course. But as I said, you still have the "stochastic parrot" problem. Now your problem is your reasoning is parroted. This is a nice problem to have, if you're just playing chess -- but is a catastrophic problem if you're hacking civil infrastructure.
If LLM can always imitate closely enough to appear as-if, how can you ever separate it from whatever actual intelligence is?
Even then, it's a pretty fragile illusion at the moment. Clearly the reasoning traces dont ground the answers. There's no intelligence taking place even as-measured by text.
But let's be clear these were always, and are, bad measures of intelligence. You cannot test a dolphin this way. And its easy to cheat on tests either thru recall , wrote-learning, etc. and IQ tests haev very poor individual test-retest reliability.
In humans there's a convenient correlation that verbal articulation in text is a strong but weak correlate of intelligence. Its "Good enough" for allocating meat bodies to our various institutions. But if you've met many well-tested people you'll realise how, in practice, terrible this measuring approach is. The world we inhabit is filled with misclassified "intellects" who perform well under text-based rubrics. Add LLMs to that heap, the cheater par execellence.
Consider thought that all mammals have imagination, and model-based reinformcement, and a wide vareity of other capacities required for intelligence. And so merely issuing "text" captures, incorrelate, only these capacities by proxy.
I'm sure if you thought about it yhou could come up with tests that distinguish lizards from birds and the greater apes from the lesser. Those are the tests