> You can make any modern LLM explain its reasoning
You can make any modern LLM create a plausible, self-consistent explanation that looks like reasoning, but it's not "the reasoning it used to arrive at that answer".
> You can make any modern LLM explain its reasoning
You can make any modern LLM create a plausible, self-consistent explanation that looks like reasoning, but it's not "the reasoning it used to arrive at that answer".
We often make a decision based on a gut feeling, and then backfill a logical reason supporting our feeling, without even realizing we're doing it -- rationalization.
When you ask people who are rationalising poor behaviour about the scenario, but it is someone else doing it, they may arrive at a better answer. Can we use multiple LLMs to achieve self criticism and critical thinking?
To your point, I agree that nominally there should be a way to give conceptual names to paths of weights, and when answering a question, notice which weights were and were not applied and retrospect on that.
That's not what reasoning traces as they currently exist are, though.
Is it unreasonable to hold LLMs to a higher standard?
I disagree. In fact, I find people who think of themselves as highly rational seem to be particularly prone to rationalization. Because being rational is integral to their whole identity, they are much more eager to accept their own rationalizations as sound and resist admitting that they are guided by instinct.
Like asking a human "how did you catch that fast ball coming at you?"
They use the "Clock" algorithm, the "Pizza" algorithm, a few other ones.
> All networks we study implement the same simple neuron model in their first-layer MLPs: degree-1 sinusoidal fits in layer 1, with deeper layers combining into degree-2 sinusoidal interactions.
https://neurips.cc/virtual/2025/loc/san-diego/133808
https://arxiv.org/abs/2502.00873
When you ask them they don't mention these at all, they give you high-school math:
https://chatgpt.com/share/6a82afdd-872c-83eb-aad7-622d27f2dc...
> did you use the "cos" or "sin" function at all during this addition computation?
> No. There is no need for trigonometric functions like sin or cos. The computation only uses basic arithmetic and place-value reasoning. Of course, if someone were implementing arithmetic in a computer, it is theoretically possible to express addition using extremely complicated formulas involving sin and cos. But in the reasoning I described, no trigonometric functions were involved at all. I simply decomposed the numbers into hundreds and smaller parts and added them.