More RLVR.
Give it verifiable problems, if it doesn't find a solution move on, if it does, use that as a reward signal.
It is entirely possible to run multiple concurrent post-training runs. When a frontier lap deploys a 1M RL gym rollout, these 1M environments are absolutely not talking to each other or interconnected. They individually generate traces and movements that can be then combined for post training.