This sort of thing imo is similar to what openAI did with transformer architecture i.e. google invented it but couldn't scale it in the right direction and deepmind got busy with atari games. They had all the pieces still openai could do it. It seems to be it comes down to research leadership in what methods to choose to invest in. But yeah, the budgets big labs have, they can easily try 10 different techniques and brute force it all but seems like they are too opinionated in methods and less urgent on outcomes.
[paper] https://arxiv.org/pdf/2501.12948 [tulu] https://x.com/hamishivi/status/1881394117810500004
Also https://epoch.ai/gradient-updates/how-has-deepseek-improved-... has a summary of all the architectural improvements DeepSeek made to increase performance.
My hypothesis is that everyone was just so stunned by oai's result so most just decided to blindly chase it and do what oai did (i.e. scaling up). And it's only after o1 people started seriously trying other ideas.
Also on the note of RL optimizers, if anyone here is familiar with this space can they comment on how the recently introduced PRIME [1] compares to PPO directly? Their description is confusing since the "implicit PRM" they introduce which is trained alongside the policy network seems no different from the value network of PPO.