Ok that makes a lot of sense.
Why do they call it reinforcement learning then? Is it not traditional RE such as Q learning?
Why do they call it reinforcement learning then? Is it not traditional RE such as Q learning?
The benefit of RL in general is that you're training on states the agent is likely to find itself in, and the cost is needing an agent which explores salient states. Which is why we keep seeing RL as a finishing step after imitation (eg AlphaStar first learning StarCraft from replays)
The LLM is trained to increase this reward score (or minimize the inverse), which is what makes it RL.