The paper is very light on details and doesn't offer any evidence of the stated hypothesis beyond referencing other work.
this position paper is actually bang on a direction we've been working on for the past year — scaling many specialized agents together with RL instead of just scaling one big model.