> how do you argue that these models are not able to reason?
I don't make this argument. Benchmarks like CLUTRR[1] show how poorly LLMs do in reasoning.
I don't make this argument. Benchmarks like CLUTRR[1] show how poorly LLMs do in reasoning.
Reasoning in general is not a binary or global property. You aren't surprised when high-schoolers don't, after having learned how to draw 2D shapes, immediately go on to draw 200D hypercubes.
The problem was never "my llm can't do addition" - it can write python code!
The problem is "my llm can't solve hard problems that require reasoning"