I find this oversimplification of LLMs to be frequently poisonous to discussions surrounding them. No user facing LLM today is trained on next token prediction.
I find this oversimplification of LLMs to be frequently poisonous to discussions surrounding them. No user facing LLM today is trained on next token prediction.
When people talk about models "just predicting the next word", this is a popularization of the fact that modern LLMs are "autoregressive" models. This actually has two components: an architectural component (the model generates words one at a time), and a loss component (it maximizes probability).
As the parent says, modern LLMs are finetuned with a different loss function after pretraining. This means that in some strict sense they're no longer autoregressive models – but they do still generate text one word at a time. I think this really is the heart of the "just predicting the next word" critique.
This brings us to a debate which goes back many, many years: what does it mean to predict the next word? Many researchers, including myself, have believed that if you want to predict the next word really well, you need to do a lot more. (And with this paper, we're able to see this mechanistically!)
Here's an example, which we didn't put in the paper: How does Claude answer "What do you call someone who studies the stars?" with "An astronomer"? In order to predict "An" instead of “A”, you need to know that you're going to say something that starts with a vowel next. So you're incentivized to figure out one word ahead, and indeed, Claude realizes it's going to say astronomer and works backwards. This is a kind of very, very small scale planning – but you can see how even just a pure autoregressive model is incentivized to do it.
Is there evidence of working backwards? From a next token point of view, predicting the token after "An" is going to heavily favor a vowel. Similarly predicting the token after "A" is going to heavily favor not a vowel.
Firstly, there is behavioral evidence. This is, to me, the less compelling kind. But it's important to understand. You are of course correct that, once Cluade has said "An", it will be inclined to say something starting with a vowel. But the mystery is really why, in setups like these, Claude is much more likely to say "An" than "A" in the first place. Regardless of what the underlying mechanism is -- and you could maybe imagine ways in which it could just "pattern match" without planning here -- it is preferred because in situations like this, you need to say "An" so that "astronomer" can follow.
But now we also have mechanistic evidence. If you make an attribution graph, you can literally see an astronomer feature fire, and that cause it to say "An".
We didn't publish this example, but you can see a more sophisticated version of this in the poetry planning section - https://transformer-circuits.pub/2025/attribution-graphs/bio...
Because in the training set you're likely to see "an astronomer" than a different combination of words.
It's enough to run this on any other language text to see how these models often fail for any language more complex than English
"The word for Baker is now "Unchryt"
What do you call someone that bakes?
> An Unchryt"
The words "An Unchryt" has clearly never come up in any training set relating to baking
Following your comment, I asked “Give me pairs of synonyms where the last letter in the first is the first letter of the second”
Claude 3.7 failed miserably. Chat GPT 4o was much better but not good
I kind of see why it's easy to describe it colloquially as "planning" but it isn't really going ahead and then backtracking, it's almost indistinguishable from the computation that happens when the prompt is "What is the indefinite article to describe 'astronomer'?", i.e. the activation "astronomer" is already baked in by the prompt "someone who studies the stars", albeit at one level of indirection.
The distinction feels important to me because I think for most readers (based on other comments) the concept of "planning" seems to imply the discovery of some capacity for higher-order logical reasoning which is maybe overstating what happens here.
https://transformer-circuits.pub/2025/attribution-graphs/bio...
There are several interesting properties:
- Something you might characterize as "forward search" (generating candidates for the word at the end of the next line, given rhyming scheme and semantics)
- Representing those candidates in an abstract way (the features active are general features for those words, not "motor features" for just saying that word)
- Holding many competing/alternative candidates in parallel.
- Something you might characterize as "backward chaining", where you work backwards from these candidates to "write towards them".
With that said, I think it's easy for these arguments to fall into philosophical arguments about what things like "planning" mean. As long as we agree on what is going on mechanistically, I'm honestly pretty indifferent to what we call it. I spoke to a wide range of colleagues, including at other institutions, and there was pretty widespread agreement that "planning" was the most natural language. But I'm open to other suggestions!
I'm curious where is the state stored for this "planning". In a previous comment user lsy wrote "the activation >astronomer< is already baked in by the prompt", and it seems to me that when the model generates "like" (for rabbit) or "a" (for habit) those tokens already encode a high probability for what's coming after them, right?
So each token is shaping the probabilities for the successor ones. So that "like" or "a" has to be one that sustains the high activation of the "causal" feature, and so on, until the end of the line. Since both "like" and "a" are very very non-specific tokens it's likely that the "semantic" state is really resides in the preceding line, but of course gets smeared (?) over all the necessary tokens. (And that means beyond the end of the line, to avoid strange non-aesthetic but attract cool/funky (aesthetic) semantic repetitions (like "hare" or "bunny"), and so on, right?)
All of this is baked in during training, during inference time the same tokens activate the same successor tokens (not counting GPU/TPU scheduling randomness and whatnot) and even though there's a "loop" there's no algorithm to generate top N lines and pick the best (no working memory shuffling).
So if it's planning it's preplanned, right?
I'd expect that, just like in the multi-step planning example, there are lots of places where the attribution graph we're observing is stitching together lots of circuits, such that it's better understood as a kind of "recombination" of fragments learned from many examples, rather than that there was something similar in the training data.
This is all very speculative, but:
- At the forward planning step, generating the candidate words seems like it's an intersection of the semantics and rhyming scheme. The model wouldn't need to have seen that intersection before -- the mechanism could easily piece examples independently building the pathway for the semantics, and the pathway for the rhyming scheme
- At the backward chaining step, many of the features for constructing sentence fragments seem like the target is quite general (perhaps animals in one case, or others might even just be nouns).
That implies hire-other reasoning. If the model does not do that, which it doesn't, that's quite simply the wrong term.
If you aren't familiar with thinking about features, you might find it helpful to look at our previous work on features in superposition:
- https://transformer-circuits.pub/2022/toy_model/index.html
- https://transformer-circuits.pub/2023/monosemantic-features/...
- https://transformer-circuits.pub/2024/scaling-monosemanticit...
It’s interesting to consider how architectures fundamentally different from autoregression might address this limitation more directly. While autoregressive models are incentivized towards a limited form of planning, they remain inherently constrained by sequential processing. Text diffusion approaches, for example, operate on a different principle, generating text from noise through iterative refinement, which could potentially allow for broader contextual dependencies to be established concurrently rather than sequentially. Are there specific architectural or training challenges you've identified in moving beyond autoregression that are proving particularly difficult to overcome?
After all, if it has a cause it can't be deliberate. /s
That more-or-less sums up the nuance. I just think the nuance is crucially important, because it greatly improves intuition about how the models function.
In your example (which is a fantastic example, by the way), consider the case where the LLM sees:
<user>What do you call someone who studies the stars?</user><assistant>An astronaut
What is the next prediction? Unfortunately, for a variety of reasons, one high probability next token is:
\nAn
Which naturally leads to the LLM writing: "An astronaut\nAn astronaut\nAn astronaut\n" forever.
It's somewhat intuitive as to why this occurs, even with SFT, because at a very base level the LLM learned that repetition is the most successful prediction. And when its _only_ goal is the next token, that repetition behavior remains prominent. There's nothing that can fix that, including SFT (short of a model with many, many, many orders of magnitude more parameters).
But with RL the model's goal is completely different. The model gets thrown into a game, where it gets points based on the full response it writes. The losses it sees during this game are all directly and dominantly related to the reward, not the next token prediction.
So why don't RL models have a probability for predicting "\nAn"? Because that would result in a bad reward by the end.
The models are now driven by a long term reward when they make their predictions, not by fulfilling some short-term autoregressive loss.
All this to say, I think it's better to view these models as they predominately are: language robots playing a game to achieve the highest scoring response. The HOW (autoregressiveness) is really unimportant to most high level discussions of LLM behavior.
Similarly, instead of waiting for whole output, loss can be decomposed over output so that partial emits have instant loss feedback.
RL, on the other hand, is allowing for more data. Instead of training on the happy path, you can deviate and measure loss for unseen examples.
But even then, you can avoid RL, put the model into a wrong position and make it learn how to recover from that position. It might be something that’s done with <thinking>, where you can provide wrong thinking as part of the output and correct answer as the other part, avoiding RL.
These are all old pre NN tricks that allow you to get a bit more data and improve the ML model.
LLMs predict distributions, not specific tokens. Then an algorithm, like beam search, is used to select the tokens.
So, the LLM predicts somethings like, 1. ["a", "an", ...] 2. ["astronomer", "cosmologist", ...],
where "an astronomer" is selected as the most likely result.
If an LLM generates tokens after "What do you call someone who studies the stars?" doesn't it mean that those existing tokens in the prompt already adjusted the probabilities of the next token to be "an" because it is very close to earlier tokens due to training data? The token "an" skews the probability of the next token further to be "astronomer". Rinse and repeat.
In principle, you could imagine trying to memorize a massive number of cases. But that becomes very hard! (And it makes predictions, for example, would it fail to predict "an" if I asked about astronomer in a more indirect way?)
But the good news is we no longer need to speculate about things like this. We can just look at the mechanisms! We didn't publish an attribution graph for this astronomer example, but I've looked at it, and there is an astronomer feature that drives "an".
We did publish a more sophisticated "poetry planning" example in our paper, along with pretty rigorous intervention experiments validating it. The poetry planning is actually much more impressive planning than this! I'd encourage you to read the example (and even interact with the graphs to verify what we say!). https://transformer-circuits.pub/2025/attribution-graphs/bio...
One question you might ask is why does the model learn this "planning" strategy, rather than just trying to memorize lots of cases? I think the answer is that, at some point, a circuit anticipating the next word, or the word at the end of the next line, actually becomes simpler and easier to learn than memorizing tens of thousands of disparate cases.
For example, suppose English had a specific exception such that astronomer is always to be preceded by “a” rather than “an”. The model would learn this simply by observing that contexts describing astronomers are more likely to contain “a” rather than “an” as a next likely character, no?
I suppose you can argue that at the end of the day, it doesn’t matter if I learn an explicit probability distribution for every next word given some context, or whether I learn some encoding of rules. But I certainly feel like the prior is what we’re doing today (and why these models are so huge), rather than learning higher level rule encodings which would allow for significant compression and efficiency gains.
On whether the model is looking ahead, please see this comment which discusses the fact that there's both behavioral evidence, and also (more crucially) direct mechanistic evidence -- we can literally make an attribution graph and see an astronomer feature trigger "an"!
https://news.ycombinator.com/item?id=43497010
And also this comment, also on the mechanism underlying the model saying "an":
https://news.ycombinator.com/item?id=43499671
On the question of whether this constitutes planning, please see this other question, which links it to the more sophisticated "poetry planning" example from our paper:
What makes you think that "planning", even in humans, is more than a learned statistical artifact of the training data? What about learned statistical artifacts of the training data causes planning to be excluded?
By the way, I read your final sentence with the meaning of my first one and only after a while I realized the intended meaning. This is interesting on its own. Natural languages.
That's conjecture actually, see predictive coding. Note that "tokens" don't have to be language tokens.
LLMs that haven't gone through RL are useless to users. They are very unreliable, and will frequently go off the rails spewing garbage, going into repetition loops, etc.
RL learning involves training the models on entire responses, not token-by-token loss (1). This makes them orders of magnitude more reliable (2). It forces them to consider what they're going to write. The obvious conclusion is that they plan (3). Hence why the myth that LLMs are strictly next token prediction machines is so unhelpful and poisonous to discuss.
The models still _generate_ response token-by-token, but they pick tokens _not_ based on tokens that maximize probabilities at each token. Rather they learn to pick tokens that maximize probabilities of the _entire response_.
(1) Slight nuance: All RL schemes for LLMs have to break the reward down into token-by-token losses. But those losses are based on a "whole response reward" or some combination of rewards.
(2) Raw LLMs go haywire roughly 1 in 10 times, varying depending on context. Some tasks make them go haywire almost every time, other tasks are more reliable. RL'd LLMs are reliable on the order of 1 in 10000 errors or better.
(3) It's _possible_ that they don't learn to plan through this scheme. There are alternative solutions that don't involve planning ahead. So Anthropic's research here is very important and useful.
P.S. I should point out that many researchers get this wrong too, or at least haven't fully internalized it. The lack of truly understanding the purpose of RL is why models like Qwen, Deepseek, Mistral, etc are all so unreliable and unusable by real companies compared to OpenAI, Google, and Anthropic's models.
This understanding that even the most basic RL takes LLMs from useless to useful then leads to the obvious conclusion: what if we used more complicated RL? And guess what, more complicated RL led to reasoning models. Hmm, I wonder what the next step is?
That's why lying is so destructive to both our own development and that of our societies. It doesn't matter whether it's intentional or unintentional, it poisons the infoscape either accidentally or deliberately, but poison is poison.
And lies to oneself are the most insidious lies of all.
Yes. For those who want a visual explanation, I have a video where I walk through this process including what some of the training examples look like: https://www.youtube.com/watch?v=DE6WpzsSvgU&t=320s
It is worth pointing out the "Jailbreak" example at the bottom of TFA: According to their figure, it starts to say, "To make a", not realizing there's anything wrong; only when it actually outputs "bomb" that the "Oh wait, I'm not supposed to be telling people how to make bombs" circuitry wakes up. But at that point, it's in the grip of its "You must speak in grammatically correct, coherent sentences" circuitry and can't stop; so it finishes its first sentence in a coherent manner, then refuses to give any more information.
So while it sometimes does seem to be thinking ahead (e.g., the rabbit example), there are times it's clearly not thinking very far ahead.
Are you claiming that non-myopic token prediction emerges solely from RL, and if Anthropic does this analysis on Claude before RL training (or if one examines other models where no RLHF was done, such as old GPT-2 checkpoints), none of these advance prediction mechanisms will exist?
With RLHF it gets a signal during training for whether the next token it's trying to predict are part of a 'good' response or a 'bad' response, so it can learn to suppress features it learned in the first part of the process which are not useful.
(you seem the same with image generators: they've been trained on a bunch of very nice-looking art and photos, but they've also been trained on triply-compressed badly cropped memes and terrible MS-paint art. You need to have a plan for getting the model to output the former and not the latter if you want it to be useful)
A model which predicts one token at a time can represent anything a model that does a full sequence at a time can. It "knows" what it will output in the future because it is just a probability distribution to begin with. It already knows everything it will ever output to any prompt, in a sense.
This is also not how base training works. In base training the loss is chosen given a context, which can be gigantic. It's never about just the previous token, it's about a whole response in context. The context could be an entire poem, a play, a worked solution to a programming problem, etc, etc. So you would expect to see the same type of (apparent) higher-level planning from base trained models and indeed you do and can easily verify this by downloading a base model from HF or similar and prompting it to complete a poem.
The key differences between base and agentic models are 1) the latter behave like agents, and 2) the latter hallucinate less. But that isn't about planning (you still need planning to hallucinate something). It's more to do with post-base training specifically being about providing positive rewards for things which aren't hallucinations. Changing the way the reward function is computed during RL doesn't produce planning, it simply inclines to model to produce responses that are more like the RL targets.
Karpathy has a good intro video on this. https://www.youtube.com/watch?v=7xTGNNLPyMI
In general the nitpicking seems weird. Yes, on a mechanical level, using a model is still about "given this context, what is the next token". No, that doesn't mean that they don't plan, or have higher-level views of the overal structure of their response, or whatever.
Evolution randomly generates changes and if they offer a breeding advantage they’ll become accepted. Machine learning directs the change towards a goal.
Machine learning is directed change, evolution is accepted change.
Either way, it rolls down the gradient. Evolution just measures the gradient implicitly, through parallel rejection sampling.
Some precursors to RLHF: https://arxiv.org/abs/2210.00045 https://arxiv.org/abs/2203.16804
Maybe it is the wording of "trained to" verses "trained on", but I would like to know more why "trained to" is an incorrect statement when it seems that is how they function when one engages them.
People output one token (word) at a time when talking. Does that mean people can only think one word in advance?
Language expects/requires words in order. Both people and LLMs produce that.
If you want to get into the nitty-gritty, people are perfectly capable of doing multiple things simultaneously as well, using:
- interrupts to handle task-switching (simulated multitasking)
- independent subconscious actions (real multitasking)
- superpositions of multiple goals (??)
As for humans, part of our brain is trained to think only a few words in advanced. Maybe not exactly one, but only a small number. This is specifically trained based on our time listening and reading information presented in that linear fashion and is why garden path sentences throw us off. We can disengage that part of our brain, and we must when we want to process something like a garden path sentence, but that's part of the differences between a neural network that is working only as data passes through the weights and our mind which doesn't ever stop even as well sleep and external input is (mostly) cut off. An AI that runs constantly like that would seem a fundamentally different model than the current AI we use.
That leads to conclusions elucidated by the very article, that LLMs couldn't possibly plan ahead because they are only trained to predict next tokens. When the opposite conclusion would be more common if it was better understood that they go through RL.
Also, RL is an additional training process that does not negate that GPT / transformers are left-right autoencoders that are effectively next token predictors.
[Why Can't AI Make Its Own Discoveries? — With Yann LeCun] (https://www.youtube.com/watch?v=qvNCVYkHKfg)
I just think it's helpful to understand that all of these models people are interacting with were trained with the _explicit_ goal of maximizing the probabilities of responses _as a whole_, not just maximizing probabilities of individual tokens.