Or is something on top?
Or is something on top?
The way to think about it is that backpropagation changes the parameters of a model so they get closer to some sort of desired output.
In pre-training and SFT, the parameters are changed so the model does a better job of replicating the next word in the training data, given the words it has already seen.
In RLHF, the parameters are changed so the model does a better job of outputting the response that aligns to the human's preference (see: the feedback screen in the linked article).
So how can you update weights without doing back-propagation? Or is it still back propagation but with a different metric?
Why do they call it reinforcement learning then? Is it not traditional RE such as Q learning?
The benefit of RL in general is that you're training on states the agent is likely to find itself in, and the cost is needing an agent which explores salient states. Which is why we keep seeing RL as a finishing step after imitation (eg AlphaStar first learning StarCraft from replays)
The LLM is trained to increase this reward score (or minimize the inverse), which is what makes it RL.
Think of it this way - there are an equal number of rude and polite comments online (actually probably way more rude ones).
If a model is trained on that data, how do you get it to only respond politely?
You could filter out the rude comments, but that's expensive and those rude comments may still have other helpful patterns that tech your model other stuff.
Alternatively, you could pre-train on the rude comments, but then after pre-training is done, you hire a ton of people in a low cost geo and ask them 'do you prefer comment 1 (a polite output of the pre-trained model) or comment 2 (a rude output).'
The model then 'learns' that comment 1 is better because it gets more votes, and adjusts parameter's (through backpropagation) to make comment 1 instead of comment 2
In practice, you can't control what the model outputs, so you just ask it to give you it's top N responses and the humans rank all of them, hoping you get a decent mix of rude and polite.
LLMs are just pattern identifiers and repeaters. They are trained on inherently biased training datasets of inherently biased text written by inherently biased humans. Every single step of training introduces some amount of bias to an LLM.
For the finetuning i'm using LoRA to freeze most of the layers for parameter optimization. Using PEFT from huggingface
Knowing that you will be doing further training on a provided model (even "just" extensive fine-tuning), one would want to distinguish the training done before you get your hands on it, from the training you do. An obvious word for that previous training is pre-training, which unfortunately conflicts with a term of art.
SFT and RLHF is attempting to further guide the model in terms of steerability + alignment of output.
In fact, the InstructGPT authors were worried about losing the pre-trained model's underlying probability distribution, so they try a version where it penalizes the model deviating too significantly from the original distribution (using KL). I don't remember them seeing a significant difference in performance.