LLMs have been around for years, and they're explicitly trained on the entire history of human mathematics (without which they'd be unable to do anything).
LLMs have been around for years, and they're explicitly trained on the entire history of human mathematics (without which they'd be unable to do anything).
I like the test proposed by Demis Hassabis. Something like: train an LLM on pre-Einstein physics and see if it can rediscover what Einstein did. Despite the incessant noise online, we seem to be no closer to this. If there's material evidence that we are, I'd sincerely like to hear about it!
2 years ago, a DeepMind system achieved IMO Silver. This system was not a pure LLM, nor were the problems solved end to end in language. In fact the role of the LLM was to translate the problem into lean and generate strategies that could be auto verified. Some problems took up to 3 days to compute.
1 year ago, both OpenAI and DeepMind announced internal models had achieved IMO Gold. Unlike the cobbled together Alphaproof and AlphaGeometry, these results were purely LLMs, solved end to end in language, with any lean proofs generated by the LLM afterwards. They achieved this result in the time limits alloted to humans (4.5 hours)
Today, we have publicly available models capable of solving well known and open standing conjectures. And in less than the time it took Alphaproof/Geometry to attain IMO Silver, internal models can even solve a millenium prize problem.
Comments like yours just make it plain you aren't paying any attention at all. The indication that it will continue is the fact that it has continued despite some assuring us we're at the precipice of some plateau at every single moment. You think it's because of 'breathless hype'? Only if you're blind.
OpenAI stole the work of Tristan Buckmaster to do it, and refused to explicitly deny it.
What do you think has enabled the 'staggering rate' of improvement? Scaling up? More data? It seems like, since LLMs do not solve problems from first principles and make crucial use of human work (without which they'd have nothing), once all the data has been eaten, some sort of plateau will indeed be reached. Anthropic are already reportedly buying up all of the world's second-hand books and destructively (their words) scanning them in a desperate rush for ever more training data, so I don't think this theory is unfounded.
What do you think of the Hassabis Test? Until it passes, I don't see how humans are replaced.
No they did not. Tristan (with heavy aid from LLMs btw) solved a sub-problem via a path Open AI's full solution didn't take (forced vs unforced). To say they stole his work would be silly.
And they explicitly say - “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way.”
https://www.nytimes.com/2026/09/10/science/tristan-buckmaste...
>What do you think has enabled the 'staggering rate' of improvement?
Compute and data
>It seems like, since LLMs do not solve problems from first principles and make crucial use of human work (without which they'd have nothing), once all the data has been eaten, some sort of plateau will indeed be reached.
Data used to train the models has become increasingly synthetic. The reinforcement learning getting them superhuman at math isn't tied to human data. If this is what you're betting on to make things stall, best to move on.
>Anthropic are already reportedly buying up all of the world's second-hand books
That is not what they are doing.
>What do you think of the Hassabis Test? Until it passes, I don't see how humans are replaced.
There's not enough pre-einstein text digitzed to make a fair shot at such an experiment so neat idea but kind of a waste of time. And it's kind of meaningless. We're already in the midst of AI becoming superhuman in one domain. We don't need any what ifs for something that is happening right now, in front of us. If some want to stick their hands in the sands, shouting 'la la humans are special' then there's nothing to do about that.
He says they stole it; see for example his recent interview with Numberphile. So is he silly? Or do you know better?
> There's not enough pre-einstein text digitzed to make a fair shot at such an experiment
Then why was it possible for humans to do physics, or indeed function at all, pre Einstein? I guess humans are special after all.
He can say whatever he wants. I'm not obligated to take it at face value, especially when it seems to fall apart with some scrutiny.
>Then why was it possible for humans to do physics, or indeed function at all, pre Einstein? I guess humans are special after all.
Billions of years of evolution probably has something to do with it.
I don’t think it’s a stretch to say that each side’s claim isn’t equally trustworthy.
So I have my reasons to weigh his statements more heavily. I certainly don’t take anything for granted.