896 karma · joined August 18, 2023
gemini 2.5 pro shines for 200k+ tokens
i still think my original statement is fair
> How can these spurious rewards possibly work? Can we get similar gains on other models with broken rewards?
it's because in those cases, RLVR merely elicits the reasoning strategies already contained in the model through pre-training
this paper, which uses Reasoning gym, shows that you need to train for way longer than those papers you mentioned to actually uncover novel reasoning strategies: https://arxiv.org/abs/2505.24864
does this mean that previous RL papers claiming the opposite were possibly bottlenecked by small datasets?
Given that GDM pioneered RL, that's a reasonable assumption
which sort of interviews will you support?
Curious how their Devin competitor will pan out given Devin's challenges