HNHacker News
TopNewBestAskShowJobs

danielmarkbruce

5,104 karma · joined March 27, 2020

submissionscomments
danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
A guess at what though? One guesses at truths they don't know, or events that haven't happened yet. What is the model guessing?
danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
So, this is the cause of the problem.... People take an intro to LLMs course, follow happily along, and don't realize there is more to it than the next token prediction. And those courses teach how LLMs were built in 2017-2020 maybe. Then RL got added to the mix. The current models really are very different to the models from then - everything that is now considered "post-training" isn't doing next token prediction.
danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
Assuming you are saying that RL is changing the model from doing one thing to another, yes. RL is changing the nature of the model.
danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
The discussion is basically: what is a model trying to do?

One may reasonably assert it isn't trying to do anything. But, in practice, if you give it an objective function and optimize it, the model is basically trained to "do" something. So what is it trained to "do"? During pre training it is trained to produce a distribution which is a prediction of the next token in it's training data samples. During RLVR and RLHF, it is trained to produce a distribution of tokens that will maximize a scoring function over many steps - not just the next step. The fact that it produces a distribution of potential choices for the next step doesn't mean the next step is a prediction. It's more of a "strategy" or "probabilistic path choice". The word used in RL is a "policy". It's a decent word to describe what the model is.

So, modern LLMs are trying to produce a good sequence of tokens. They are "good token sequence producer machines". Not "next token prediction machines". Pre RLHF (in practice, go back to pre chatgpt) they really were "next token prediction machines".

danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
No one is arguing about the architecture of the model. It's the objective function and optimizer.
danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
This is pedantic, but, actually RL has improved the quality of sentence construction in LLMs quite dramatically... And once you do some RL on that model, it aint a next token prediction machine any longer.
danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
Probably the easiest way to describe an LLM that it's a policy. There is a reason that word has stuck in RL.

And it's not just RLVR. RLHF has been going on for years and years. LLMs have not been "next token predictors" for probably 5-6 years.

danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
It's not an estimation of something. It's a policy.
danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
You are conflating "half built" with "a piece of a system".

The model weights change as the model goes through the training process. They aren't stored after pre-training is done and other weights are put somewhere else. It's more like pottery - the thing changes. It's not correct to say something is soft and malleable because it once was.

danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
Yes, it was. Nobody says a system or product works a certain way and means the system while it's half built. "Bridges drop cars in the water!". Right.

You aren't in this field. You are clearly wrong and just can't handle it.

danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
Your mistake was assuming people would be bothered to understand the details of how things work. Most people are lazy and don't know the details of how anything works.
danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
They aren't predicting the next token. It's quite literally not a prediction.
danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
All modern LLMs that actually get used go through post-training. The finished product is something which has been through post training. So they are not next token prediction machines.
danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
No, it's not philsophical. Because if you optimize to predict, you are doing something different to optimizing for a reward. It's a different process - different objective function, different optimization, different set up.
danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
Read through the article and comments. You are talking solely about pre-training. I'm talking about post training.

Respectfully, you are miles out of your depth. GPT-2 didn't use any reinforcement learning and is often given as a toy example. That release was 2019 and models now go through a various phases of training with different objective functions and optimizers.

danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
Nope. This isn't right.
danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
Yup, you are mostly right.

I guess the people in my camp find the "it's just a next token predictor" stupid in that it's like saying "it's just a bunch of carbon and hydrogen", but it's also one of those things where people like to think they are clever because they think they are theoretically correct. But they aren't even that. So it's like double stupid. But the "next token predictor" part is at least technically correct (like, carbon and hydrogen right) for pretraining, so the debate can't really be won there.

danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
Lol, sure, just read a blog post and you'll understand how a car works....It's very simple....
danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
Predict implies you don't control a situation. That's the difference.
danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
People have and are trying things. Lots and lots of things. They just don't go around promoting failed ideas.
danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
The fight is about the predictor language in some cases. Because it's only a trivial difference to those who don't understand the details of how these things are made. In pre-training the model really is trained to predict the next token. What is being emitted by the model is, by structure, by training and by optimization, a prediction of the very next token.

What is emitted by a model during RLHF and RLVR is not, by structure, training or optimization, a prediction of the next token.

danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
It's not a prediction of the next move though, and that is the point. It's a prediction of what will happen if you make that move.

So, it's not a next move predictor. It's a game result predictor.

danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
Emitting and predicting are different things though. Prediction implies there is some "truth" or event or something that you can test against. Prediction implies the model just learns from existing text, and optimizes to predict the next token in training data. That's just not true.
danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
If you are going to say "literally", then what is your literal definition for the word "prediction" ?
danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
If you haven't built one, and don't understand how they work, why comment?
danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
Respectfully, go build one, including doing RLHF and RLVR. Those phases generate lots of tokens, then get scored on the entirety of the output, then optimize based on a scoring of that output. It doesn't check a "prediction" against what was actually "next" in data, because there isn't any "next token" data it's training on.
danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
Nope, it doesn't.

No logic required, you can just build an LLM yourself, including post training. You'll see that predicting the next token isn't something the model does or is optimized for in RLHF or RLVR. You can hand wave all you like, but you have never done it.

danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
Post train a model, you'll be able to determine it is not.
danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
There isn't a truth to test against. If I predict the next word in a sequence is "sat", we can check against the sequence. If I predict the roll of a die will be 4, we can check against it. Whether i give 100% or give a probabilistic prediction, we can check against the truth.

If I choose a specific move in chess, it's a choice. It's not a prediction. I might get a score 40 moves later given my choice, but I'm not predicting the next move.

To compare - during pre-training, the model literally tries to predict the next token (probabilistically), the training loop checks against the "right" answer, and the weights are updated based on that check. It's optimized to predict the next token.

danielmarkbruce··on “Next-token predictor” is the wrong mental model for LLMs
There is no truth for RLHF or RLVR. You can't predict against something if you can't check against the truth.

It's not pedantry. The objective function changes. The optimization changes. THese are real things when training a model, not hand wavy philosophical ideas.

← PreviousPage 4 of 34Next →