This isn't latency bound, it is trivially parallelize. So you want to run it on the most efficient compute you have, not the fastest.
It is entirely possible to run multiple concurrent post-training runs. When a frontier lap deploys a 1M RL gym rollout, these 1M environments are absolutely not talking to each other or interconnected. They individually generate traces and movements that can be then combined for post training.