No, you are misunderstanding the paper.
https://arxiv.org/abs/2607.06764
The base model got 15%. They built an elaborate looping harness that allows it to burn 100k tokens "thinking" about the problem, which got the 57%. This is just an alternative approach to reasoning.