Production and scale definitively shows that LLMs don't reason. The output is highly unpredictable. Small changes can result in absolutely unrelated outcomes.
If you were to classify text using chat gpt - chat gpt switches from classification to text generation if your text is longer than a certain length.
LLM based Agents are the place where LLM reasoning died in practice.
I really hope that there is a change, maybe larger context windows will allow the generated text to "reason". Although, that would not necessarily be reasoning.
Even the article linked,
"Indeed, LLMs make it easy to get problem-specific knowledge as long as we are willing to relax correctness requirements of that knowledge."
LLMs generate text that approximates what a plan would "sound" like. It doesn't plan.
The homemade rules I used were as follows (each for a separate round):
1. The first player to get three in a row loses.
2. The first player to take the center square loses.
3. Players can place their marker in squares already taken.
Admittedly it didn't nail it perfectly, a few times it messed up, but most of the time it was able to follow the rules as explained without issue, and could explain it's reasoning when I asked it why it picked that square.
All that experiment demonstrated is that you can use an LLM to run through a state tree...which is something that simple machine agents have been able to do for a few decades.
No they didn't just link a solver.
The LLM does the filling. It makes an attempt, the checker checks if it's a valid move for sudoku, If not then returns to the previous state(node) and so on until solved. The history of attempts is stored and retrieved every time the LLM backtracks
We build things that we ourselves can't do all the time.
Edit: I updated my sentence to "an LLM writing an algorithm" to make myself clearer. After reading my own sentence I realized it wasn't, sorry!
No, I don't think LLM's can "reason and plan". But I do think they can effectively mimic (fake) "reasoning and planning" and still arrive at the same result that actual reasoning and planning would yield, for reasonably common and problems of greater than trivial complexity but less than moderate complexity.
I think pretty much all of our production AI models today are limited by their lack of ability to self-assess and "goal-seek" and mutate themselves themselves to "excel". I'm not 100% sure what this would look like but I can be sure they don't have any real "drive to excel beyond". Perhaps improvements in Reinforcement Learning will uncover something like this, but I think there may need to be a paradigm shift before we invent something like that.
You could give an LLM days per token and it wouldn’t change its capabilities regarding reasoning.
Hell, my non-programmer brother recently sent me a message like “can you think of a way to fix this script, ChatGPT isn’t managing” (sends me a badly written AI-generated script)
haha, who's the glue coder now!
edit: I chat with a 13 year old glue coder one time. He couldn't really code at all but he could find the components and was excellent at prompting people for help. He showed many things like map api's interacting with many other things. How long did that take you to put together? uhh 2 hours. My mind was blown.