In a sense, how do you know whether the problem your AI has just solved is important? A simple proxy is to just check whether humans have thought it's important.
That's also why famous open problems are a good benchmark or proxy: you don't need to convince the rest of the world that the problem your lab's new AI just solved is actually useful or hard.
There's also a difference between important and hard. There are important problems that turn out to be easy, and hard problems that turn out to be useless.
As usual, I think everyone would agree that something like "curing cancer" would be both hard and important!
I think the AI labs' current obsession with showing off mathematical results to uninformed outsiders is a cheap trick. If they were really interested in 'enabling human flourishing' (rather than just wowing, by any means possible, investors with more money than sense), they'd be showing off cures to diseases rather than solving obscure problems in combinatorics only previously considered by three Russians fifty years ago and then declaring that The Singularity is here.
LLMs have been around for years, and they're explicitly trained on the entire history of human mathematics (without which they'd be unable to do anything).
I like the test proposed by Demis Hassabis. Something like: train an LLM on pre-Einstein physics and see if it can rediscover what Einstein did. Despite the incessant noise online, we seem to be no closer to this. If there's material evidence that we are, I'd sincerely like to hear about it!
2 years ago, a DeepMind system achieved IMO Silver. This system was not a pure LLM, nor were the problems solved end to end in language. In fact the role of the LLM was to translate the problem into lean and generate strategies that could be auto verified. Some problems took up to 3 days to compute.
1 year ago, both OpenAI and DeepMind announced internal models had achieved IMO Gold. Unlike the cobbled together Alphaproof and AlphaGeometry, these results were purely LLMs, solved end to end in language, with any lean proofs generated by the LLM afterwards. They achieved this result in the time limits alloted to humans (4.5 hours)
Today, we have publicly available models capable of solving well known and open standing conjectures. And in less than the time it took Alphaproof/Geometry to attain IMO Silver, internal models can even solve a millenium prize problem.
Comments like yours just make it plain you aren't paying any attention at all. The indication that it will continue is the fact that it has continued despite some assuring us we're at the precipice of some plateau at every single moment. You think it's because of 'breathless hype'? Only if you're blind.
OpenAI stole the work of Tristan Buckmaster to do it, and refused to explicitly deny it.
What do you think has enabled the 'staggering rate' of improvement? Scaling up? More data? It seems like, since LLMs do not solve problems from first principles and make crucial use of human work (without which they'd have nothing), once all the data has been eaten, some sort of plateau will indeed be reached. Anthropic are already reportedly buying up all of the world's second-hand books and destructively (their words) scanning them in a desperate rush for ever more training data, so I don't think this theory is unfounded.
What do you think of the Hassabis Test? Until it passes, I don't see how humans are replaced.
No they did not. Tristan (with heavy aid from LLMs btw) solved a sub-problem via a path Open AI's full solution didn't take (forced vs unforced). To say they stole his work would be silly.
And they explicitly say - “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way.”
https://www.nytimes.com/2026/09/10/science/tristan-buckmaste...
>What do you think has enabled the 'staggering rate' of improvement?
Compute and data
>It seems like, since LLMs do not solve problems from first principles and make crucial use of human work (without which they'd have nothing), once all the data has been eaten, some sort of plateau will indeed be reached.
Data used to train the models has become increasingly synthetic. The reinforcement learning getting them superhuman at math isn't tied to human data. If this is what you're betting on to make things stall, best to move on.
>Anthropic are already reportedly buying up all of the world's second-hand books
That is not what they are doing.
>What do you think of the Hassabis Test? Until it passes, I don't see how humans are replaced.
There's not enough pre-einstein text digitzed to make a fair shot at such an experiment so neat idea but kind of a waste of time. And it's kind of meaningless. We're already in the midst of AI becoming superhuman in one domain. We don't need any what ifs for something that is happening right now, in front of us. If some want to stick their hands in the sands, shouting 'la la humans are special' then there's nothing to do about that.
He says they stole it; see for example his recent interview with Numberphile. So is he silly? Or do you know better?
> There's not enough pre-einstein text digitzed to make a fair shot at such an experiment
Then why was it possible for humans to do physics, or indeed function at all, pre Einstein? I guess humans are special after all.
I don’t think it’s a stretch to say that each side’s claim isn’t equally trustworthy.
So I have my reasons to weigh his statements more heavily. I certainly don’t take anything for granted.