Can’t this be extended quite far? Use a cerebras-served model, use verification techniques to generate and solve millions of problems and then use that as training?
It is entirely possible to run multiple concurrent post-training runs. When a frontier lap deploys a 1M RL gym rollout, these 1M environments are absolutely not talking to each other or interconnected. They individually generate traces and movements that can be then combined for post training.