Reinforcement Learning: An Overview
arxiv.org
arxiv.org
I struggled a lot with the first chapter, and had to look up a lot of terms that weren’t defined. But ultimately it was one of the most worthwhile things I read, and has helped me follow along with other important papers.
Did you mean section 5.6? That's LLMs and RL. Section 5.4 is Imitation Learning.
The term is also used/linked in the fourth paragraph of the Wikipedia page for Reinforcement Learning.
It’s much more a table-stakes for talking about what problems the field tries to solve term than an exclusive preserve of the deeply immersed term.
It’s a bit like a primer on machine learning using the word “regression” casually a few times before actually defining it. Good editing practice? No. An actual road block to learning? Also no.
RLHF Book
I‘ve heard people say that GRPO giving a zero gradient in cases where neither the current sample nor the group scores give any reward is advantageous for optimization. It avoids killing your base model with low signal to noise updates, that can be a problem in PPO where the critic usually causes a non-zero gradient even for samples where one would rather be like „problem too hard for now, skip.“
I‘d be curious to hear you lay out your thoughts though!
Flipping it around, if you swapped out the neural reward model in PPO with a reward function that can return zero, I thiiinnkkk it would be able to produce zero (or very low) gradient updates.
I'll be the first to admit that I don't know enough about the space to say though. I'm still a beginner here.