I wonder whether they implemented the GRPO correction from this paper, which fixes overly long response lengths: https://arxiv.org/abs/2503.20783
I guess probably not, as they don't mention it.
I guess probably not, as they don't mention it.
Interestingly, this is the same bug that most open-source LLM training frameworks (such as HF Trainer) had and only recently fixed.
In short, I'm working on a quick fix, after that, using sum or mean should yield equivalent results.
P.S. Fixed!