I am skeptical. Intuitively I don't see what self-play achieves beyond straight RL. Have the authors done a comparison with the performance they can get by RL finetuning a single model by itself?
Also this style of tasks is prone to overfitting. i.e. instead of predicting, the model just memorises what the results are.