This paper is over 1 year old and didn't review the literature at the time well enough to see the early work that led to current generation reasoning models.
Now it's out of date, automated RL on math problems seems to work and scales with compute. As we scale available compute 100x over the next 5 years and reduce cost of compute by around 10x over the same time frame, it will become increasing clear that LLMs running for a long time are capable of replacing most mathematics research.