It doesn't seem clear whatsoever that this is true? Is there evidence that LLMs are very skilled at generalizing across domains of mathematics where the training distribution sees little overlap?
As far as I can tell, this is a victory for verifiable loops using LEAN, reinforcement learning, and oodles of compute. I haven't seen evidence yet that this is proof of broad generalization beyond the training distribution.