LLMs that haven't gone through RL are useless to users. They are very unreliable, and will frequently go off the rails spewing garbage, going into repetition loops, etc.
RL learning involves training the models on entire responses, not token-by-token loss (1). This makes them orders of magnitude more reliable (2). It forces them to consider what they're going to write. The obvious conclusion is that they plan (3). Hence why the myth that LLMs are strictly next token prediction machines is so unhelpful and poisonous to discuss.
The models still _generate_ response token-by-token, but they pick tokens _not_ based on tokens that maximize probabilities at each token. Rather they learn to pick tokens that maximize probabilities of the _entire response_.
(1) Slight nuance: All RL schemes for LLMs have to break the reward down into token-by-token losses. But those losses are based on a "whole response reward" or some combination of rewards.
(2) Raw LLMs go haywire roughly 1 in 10 times, varying depending on context. Some tasks make them go haywire almost every time, other tasks are more reliable. RL'd LLMs are reliable on the order of 1 in 10000 errors or better.
(3) It's _possible_ that they don't learn to plan through this scheme. There are alternative solutions that don't involve planning ahead. So Anthropic's research here is very important and useful.
P.S. I should point out that many researchers get this wrong too, or at least haven't fully internalized it. The lack of truly understanding the purpose of RL is why models like Qwen, Deepseek, Mistral, etc are all so unreliable and unusable by real companies compared to OpenAI, Google, and Anthropic's models.
This understanding that even the most basic RL takes LLMs from useless to useful then leads to the obvious conclusion: what if we used more complicated RL? And guess what, more complicated RL led to reasoning models. Hmm, I wonder what the next step is?
That's why lying is so destructive to both our own development and that of our societies. It doesn't matter whether it's intentional or unintentional, it poisons the infoscape either accidentally or deliberately, but poison is poison.
And lies to oneself are the most insidious lies of all.
Yes. For those who want a visual explanation, I have a video where I walk through this process including what some of the training examples look like: https://www.youtube.com/watch?v=DE6WpzsSvgU&t=320s
It is worth pointing out the "Jailbreak" example at the bottom of TFA: According to their figure, it starts to say, "To make a", not realizing there's anything wrong; only when it actually outputs "bomb" that the "Oh wait, I'm not supposed to be telling people how to make bombs" circuitry wakes up. But at that point, it's in the grip of its "You must speak in grammatically correct, coherent sentences" circuitry and can't stop; so it finishes its first sentence in a coherent manner, then refuses to give any more information.
So while it sometimes does seem to be thinking ahead (e.g., the rabbit example), there are times it's clearly not thinking very far ahead.
Are you claiming that non-myopic token prediction emerges solely from RL, and if Anthropic does this analysis on Claude before RL training (or if one examines other models where no RLHF was done, such as old GPT-2 checkpoints), none of these advance prediction mechanisms will exist?
With RLHF it gets a signal during training for whether the next token it's trying to predict are part of a 'good' response or a 'bad' response, so it can learn to suppress features it learned in the first part of the process which are not useful.
(you seem the same with image generators: they've been trained on a bunch of very nice-looking art and photos, but they've also been trained on triply-compressed badly cropped memes and terrible MS-paint art. You need to have a plan for getting the model to output the former and not the latter if you want it to be useful)
A model which predicts one token at a time can represent anything a model that does a full sequence at a time can. It "knows" what it will output in the future because it is just a probability distribution to begin with. It already knows everything it will ever output to any prompt, in a sense.
This is also not how base training works. In base training the loss is chosen given a context, which can be gigantic. It's never about just the previous token, it's about a whole response in context. The context could be an entire poem, a play, a worked solution to a programming problem, etc, etc. So you would expect to see the same type of (apparent) higher-level planning from base trained models and indeed you do and can easily verify this by downloading a base model from HF or similar and prompting it to complete a poem.
The key differences between base and agentic models are 1) the latter behave like agents, and 2) the latter hallucinate less. But that isn't about planning (you still need planning to hallucinate something). It's more to do with post-base training specifically being about providing positive rewards for things which aren't hallucinations. Changing the way the reward function is computed during RL doesn't produce planning, it simply inclines to model to produce responses that are more like the RL targets.
Karpathy has a good intro video on this. https://www.youtube.com/watch?v=7xTGNNLPyMI
In general the nitpicking seems weird. Yes, on a mechanical level, using a model is still about "given this context, what is the next token". No, that doesn't mean that they don't plan, or have higher-level views of the overal structure of their response, or whatever.