This assumes that there aren't algorithmic breakthroughs which reduce training/inference costs by several OOMs.
How much do these models need to do before people throw their hands in the air and say, ok this is happening. The Erdos unit distance problem, which as far as I understand was approached by multiple competent mathematicians was solved by a frontier model. Sure people argue there was no novelty there (I cannot comment as a non-mathematician) but it feels like they can draw lines laterally from deep knowledge in different fields (in this case combinatorics and algebraic number theory I believe) and solve problems.
Now if you have millions of instances running in parallel, all "probabilistic", working on frontier AI research I really don't see the blocker (and believe me I wish I did).