The "Chain of Thought" Delusions
twitter.com
twitter.com
> Ever since I came across the Chain of Though (CoT) for LLMs paper, I wondered how it can possibly make sense given that there is little reason to believe that LLMs can follow procedures and unroll them for the current problem at hand. (After all, if they can do that, they should be able to do verification and general reasoning too--and we know they suck at those).
I don't follow. Why would procedure following imply general reasoning?
Probably he's better tuned in to the tone of the discourse than me but I feel like if you showed CoT helped with Blocksworld then that would be a huge and triumphant result, rather than something taken for granted that we must loudly disprove!
>> And yet the practitioners of CoT swear that any and every problem can be solved with LLM by giving it a bit of a CoT help.
For example, see this arxiv paper:
Generalized Planning in PDDL Domains with Pretrained Large Language Models
https://arxiv.org/abs/2305.11014
Where the authors conclude:
In this work, we showed that GPT-4 with CoT summarization and automated debugging is a surprisingly strong generalized planner in PDDL domains.
The author of the tweet is an expert on planning and he's responding to that kind of thing.
> Where the authors conclude:
> ...is a surprisingly strong generalized planner
Wtf is going on here?
What he says is that they can solve blocks world problems if you take them by the hand and show them how to do it, for every single problem you wan them to solve. Which is not very useful at all. He's arguing that this is what CoT achieves: it takes the LLM by the hand and introduces domain knowledge that the user has in every step of the way, so that it's not the LLM that's solving anything but the person using the LLM with CoT.
Do read the linked paper, it goes over this in more detail. It's a bit unfortunate that Rao chooses to communicate through the very noisy medium of twitter, but that's the internet.
The important step is that they demonstrate that there are no architectural limits that keep the LLM from acting in these domains, "only" knowledge/planning ones. Once we have a big enough dataset of CoT prompts, the model "just" has to generalize CoT prompting, not CoT following. It decomposes the problem into two halves, and demonstrates that if one half (instruction generation) is provided, the other (instruction following) becomes very tractable.
As to blocks world planning in particular, LLMs already have plenty of examples of block stacking problems - those are the standard motivational experiment in planning papers, like solving mazes is for Reinforcement Learning. Google returns 10 pages of results for "blocks world planning". If LLMs were capable of generalising as well as you expect they will one day, they should already be capable of solving block stacking problems without CoT and with no guidance.
I think mostly, LLMs at the moment are incredibly uneven. LLM assistants can pull obscure knowledge out of nowhere one second and fail extremely basic reasoning the next. So just because some example happens a lot in the source material doesn't mean the LLM can learn it. IMO, that CoT works at all is more down to luck of the training set than any inherent capability of the LLM.
We need a good search strategy that fits LLMs before we can get AGI. Maybe Chain Of Thought can become such a strategy, but it always felt too clunky to me.
Beam search already exists, it precisely allows the model to backtrack on its past tokens and find the global maxima of probability, at least within set depth/breadth. I think this is just a larger limitation of their shallow world model. Perhaps giving transformers registers will remedy this.
The thing to keep in mind in all these discussions about LLM planning and reasoning is that, when experts on planning and reasoning say that LLMs can't do planning and reasoning they have a very different, formal definition of those things in mind, than everybody else who hasn't studied them. Even some of the research papers on planning and reasoning with LLMs play very fast and loose with the use of those words and that's probably why there are so many positive results that later prove to be duds.
Now is probably a good time to re-educate computer scientists and AI researchers about planning and reasoning.
Wikipedia has a very short introduction to STRIPS planning:
https://en.wikipedia.org/wiki/Stanford_Research_Institute_Pr...
And here's a more comprhensive, but still short, introduction to automated planning and scheduling in general:
Have they tried that?
There are big problems with hallucinations because LLMs are not smart enough to know when they're starting to make mistakes.
But there's lots of work in this area, and generally in different ways to nail neural and symbolic systems together.
Planning and reasoning have long standing meaning in AI, based on explicit knowledge representation and inference. It's a subspace of AI that predates LLMs by decades.
But LLMs don't do that kind of planning and reasoning, and CoT is very clearly not a reasoning technique in the AI sense.
But frankly this post reads like someone that doesn't understand much about CoT in the first place, let alone the various methods that have improved upon it since.
It reads like one of the ad nauseum "look at me use this tool poorly, clearly it's a poor tool" examples.
In general, I've noticed mathematicians, computer scientists, and engineers tend to be very poor at evaluating LLMs because they just aren't very good at correctly identifying the scope and depth of what was modeled in the training data in the first place.
It's getting boring watching people foolishly try to evaluate things like "stack these clear blocks" (because that's something I regularly saw in social media posts) while glossing over or actively sabotaging the unbelievable modeling/simulating of much more complex and higher order critical reasoning behind various applications of things like empathy or psychological modeling.
For anyone reading this who wants to have a fun project, try creating two versions of a set of word puzzles with the same underlying logic structure. One where the problem and solution are using engineering-ish language like "clear blocks" and another using emotional/social language like "grieving friends."
Even as we wait for the next generation of models, there's a lot of people criminally underestimating and underutilizing the current models because they can't look beyond their own specialized domain languages.
If an Internet posts explains its reasoning, step-by-step it reaches a correct conclusion more often than other posts. Therefore LLMs that explain their reasoning and also more likely to reach the correct solution.
Therefore LLMs will suck at verification and general reasoning until we refine or augment our datasets.
Having said that, this has been covered and answered many times. Can we please create a responder for any HN question that can be answered with "LLMs are Internet simulators, and they do that because that is how Internet posts are."
(That's also why we can explain answers backwards as easily as forwards.)
If it works…then the why is somewhat secondary though still interesting
As you can see, the ease of giving CoT advice worsens drastically as we go from domain independent to domain specific to goal class specific to lexicographic goal-specific.
So he's arguing that CoT needs a lot of work from the user and it can only solve easy problems anyway.
Treat it as a text-based addon to your brain, IMO. A human brain has many components necessary for general intelligence. Some of those components can now be augmented by a giant neural net; others still benefit from human execution. All the LLM programming work I do is based on a cooperative dialog, where I fill in the gaps where the LLM doesn't have an appropriate pattern. Which are many. (But getting less.)
At any rate, the idea that LLMs regurgitate their training data is completely unrelated to this and afaict one of those "true but trivial, or important but false" ones. (True in the sense that humans also represent a reflection of their inputs; false in the sense that what's going on is definitely more than a collage of samples.)