Can LLMs Reason and Plan?
cacm.acm.org
cacm.acm.org
https://tidybot.cs.princeton.edu/ https://innermonologue.github.io/
Anyone who wants LLMs to plan and is actually interested in teasing the extent of those abilities knows how to structure planning requests.
It's extra funny because humans can't actually generate plans the way he tests LLMs to either.
Also Seeing output from GPT that demonstrates intelligence, reasoning, or whatever, and saying it is not real reasoning/Intelligence etc, is like looking at a plane soar and saying that the plane is fake flying. And this isn't a nature versus artificial thing either. The origin point is entirely arbitrary.
You could just as easily move the origin to Bees and say, "oh, birds aren't really flying". You could move it to planes and say, "oh, helicopters aren't really flying." It's a very meaningless statement.
Internal processes are entirely irrelevant.
If one really wants to know if LLMs "are capable of planning", one should keep an open mind for how planning behavior can manifest in textual form, then actually try to find any manifestations. Imposing one's view of how planning should look like is bad science. All they've proved is that their task format sucks for current generation LLMs.
Hell, my non-programmer brother recently sent me a message like “can you think of a way to fix this script, ChatGPT isn’t managing” (sends me a badly written AI-generated script)
haha, who's the glue coder now!
edit: I chat with a 13 year old glue coder one time. He couldn't really code at all but he could find the components and was excellent at prompting people for help. He showed many things like map api's interacting with many other things. How long did that take you to put together? uhh 2 hours. My mind was blown.
We build things that we ourselves can't do all the time.
Edit: I updated my sentence to "an LLM writing an algorithm" to make myself clearer. After reading my own sentence I realized it wasn't, sorry!
No, I don't think LLM's can "reason and plan". But I do think they can effectively mimic (fake) "reasoning and planning" and still arrive at the same result that actual reasoning and planning would yield, for reasonably common and problems of greater than trivial complexity but less than moderate complexity.
I think pretty much all of our production AI models today are limited by their lack of ability to self-assess and "goal-seek" and mutate themselves themselves to "excel". I'm not 100% sure what this would look like but I can be sure they don't have any real "drive to excel beyond". Perhaps improvements in Reinforcement Learning will uncover something like this, but I think there may need to be a paradigm shift before we invent something like that.
The homemade rules I used were as follows (each for a separate round):
1. The first player to get three in a row loses.
2. The first player to take the center square loses.
3. Players can place their marker in squares already taken.
Admittedly it didn't nail it perfectly, a few times it messed up, but most of the time it was able to follow the rules as explained without issue, and could explain it's reasoning when I asked it why it picked that square.
Production and scale definitively shows that LLMs don't reason. The output is highly unpredictable. Small changes can result in absolutely unrelated outcomes.
If you were to classify text using chat gpt - chat gpt switches from classification to text generation if your text is longer than a certain length.
LLM based Agents are the place where LLM reasoning died in practice.
I really hope that there is a change, maybe larger context windows will allow the generated text to "reason". Although, that would not necessarily be reasoning.
Even the article linked,
"Indeed, LLMs make it easy to get problem-specific knowledge as long as we are willing to relax correctness requirements of that knowledge."
LLMs generate text that approximates what a plan would "sound" like. It doesn't plan.
You could give an LLM days per token and it wouldn’t change its capabilities regarding reasoning.
All that experiment demonstrated is that you can use an LLM to run through a state tree...which is something that simple machine agents have been able to do for a few decades.
No they didn't just link a solver.
The LLM does the filling. It makes an attempt, the checker checks if it's a valid move for sudoku, If not then returns to the previous state(node) and so on until solved. The history of attempts is stored and retrieved every time the LLM backtracks
When I ask GPT4 to write some code to perform a specific task, and it does so, and the code works and correctly performs the task, is that bullshit?
It's true that LLMs are not reliable enough to trust blindly, the same applies to humans.
This is a large and dubious assumption. Most of what humans do require step coordination, but almost none of it is a result of explicit planning or verbalization.
I would also have welcomed a comparison to humans. If we apply this test to 100 humans, can we conclude humans don't reason if only 30 get it right?
Much of the criticism and skepticism around LLMs rests on a double-standard that itself rests on an almost embarrassing lack of understanding of how humans themselves operate.
They are a fascinating augmentation to existing capabilities that fills an ability to operate and do some approximation to reason and plan in an abstract semantic space, using natural language in a semantic way, to perform an abductive style of “reasoning” that has alluded us to date. Need to play chess? Use a chess playing system. But you can glue together special purpose systems for planning, reasoning, optimizing, calculating, etc, with an LLM in ways that are much more flexible and adaptable in a real world context than we’ve ever been able to achieve. That’s the magic of them.
The fact they can do pretty alright at some of these tasks is interesting and shows how powerful language is that embedded in its semantic structure is an awful lot of what you need to reason, plan, etc. But beyond the academic question, why would you not just use a provably optimal algorithm for a specific task? Interestingly, given a set of operations in context, LLMs are pretty good at identifying a situation where an operation is applicable and making the API call, and as we learn how to embed them, they’ll be able to defer to and be deferred to more accurately.
We're not even close to this result yet but LLMs seem the closest.
We may need another generation of LLMs that are more competent.
The prompt format made absolutely no sense considering they had decided to translate away from the PDDL to arbitrary natural language. So to avoid triggering their overused Clever Hans defense, I fed GPT 4 their prompt with only the instruction:
"Think critically about how we could represent these rules in a way that's clearer to an LLM. Lean into using coding style identifiers where possible, and JSON formatting"
(Hopefully the author won't claim telling a model to generate JSON is secretly telling it how to move blocks!)In a fresh context window at 0 temp and gpt-4-0613 I entered the JSON formatted rules it generated, along with the instruction:
Return a JSON array of [{[<valid actions>],<internal thought>, <action>}] that results in goal state
... the resulting answer solved their failed few-shot example with zero-shot_
Also funny blunder: their chain of thought example generates thoughts... after the action. Surely the author understands a transformer model can't rely on an ungenerated token to affect the action taken?
Edit: Also their "disguised difficulty" version was wrong?
> To perform Attack action, the following facts need to be true: Province object, Planet object, Harmony.
The single word "Harmony" somehow replaced replaces "Your hands are empty"?
Before that they claim it's meant to be a 1:1 replacement of entities, but the actual disguised versions are not longer valid instructions
LLM's live in "the world that has been written about", not the real world, and thus cannot formulate new ideas or hypothesis other than by accident. This, coupled with the lack of an ontological system for evaluating the validity of the statements it makes about a complex system, and, its lack of causal reasoning, means they cannot effectively plan.
I've worked on research related to causality that used LLM's (admittedly, pre ChatGPT and using much smaller models) and it was not uncommon to see extremely bogus causal relationships inferred such as "rising cost of living in NYC caused a flood in Argentina".
There has been some research in applying GPT-4 to Hierarchical Task Networks (HTN), one means of doing computerized semi-automated/automated planning of a complex task as a tree of less and less complex tasks [1].
There are other types of planning. Automated planning works better as there are more defined the tasks in a plan, less ambiguity in dependencies, more separate between the tasks. The OP article touches on that, noting LLMs are good at extracting planning knowledge but not good in their experience at creating executable plans. This is why I think the hybrid approach is best, using an LLM to inform and tweak other planning tools in order to create an executable plan.