Their paper https://mistral.ai/static/research/magistral.pdf is also cool! They edited GRPO via:
1. Removed KL Divergence
2. Normalize by total length (Dr. GRPO style)
3. Minibatch normalization for advantages
4. Relaxing trust region
1. Removed KL Divergence
2. Normalize by total length (Dr. GRPO style)
3. Minibatch normalization for advantages
4. Relaxing trust region
The paper they cite "What matters in on-policy RL" claims it does not lead to much difference on their suite of test problems, and (mean-of-minibatch)-normalization doesn't seem theoretically motivated for convergence to the optimal policy?
Wait, how are they computing the loss?
The goal of it was to "force" the model not to stray to far away from the original checkpoint, but it can hinder the model from learning new things