They don’t have any more patience for this.
I even think it's plausible a lot of mathematicians are excited by it, but the sweeping confidence of the comment you replied to without anything to back it up leaves some to be desired
It looks like people are enjoying themselves, having fun with the new results, and generally doing all of the things you say "science" is supposed to be about. So what's the problem?
https://www.youtube.com/watch?v=LKiBlGDfRU8 https://www.youtube.com/watch?v=shFUDPqVmTg
It is entirely possible that one day progress just stops or slows down, but with current evidence, I don't find that too likely - at least not in the near future. The sheer amount of resources being put into this (AI) race is mind-boggling.
So while past performance does not guarantee future results, I'm just going to kick back, and assume that many of the current issues will be fixed with future models.
Maybe this is just a matter of what model developers choose to invest training resources in, but I don’t think it’s inevitable unless clarity is made a higher priority
The problem is that right now mathematicians don't have the economical incentive to read these AI generated results. Even if you love mathematics and all that, it's always more important to get a job, and for that it doesn't seem like a good idea to invest time around problems that AI touches because you can't compete with it and you don't know if tomorrow they'll improve by x10 the sota.
Of course, it's not clear at this point whether reporting such a result even matters, but still. In its own right, it's a very cool result.
What do you do if those results suck like in the article?
But providing the answer in gibberish along with a certificate is not that, it's at best a cruel way to do it, but I'm leaning towards the idea that it's a fundamental misunderstanding of what it means to do math and what it means to communicate a result.
If you think sending an answer in gibberish is acceptable just because it's true then SSdtIG5vdCBzdXJlIHdoYXQgdG8gdGVsbCB5b3UsIGJ1dCB3ZSBkaXNhZ3JlZSBvbiB0aGF0.
They normally don't feel like they are in some kind of race to publish the results ASAP and claim priority. Cases like that are very rare (but they get media coverage because they are so unusual).
OpenAI did a publicity stunt, their motivation is not to make a good contribution to the field, which has very different standards and culture, compared to the AI labs.
As the author of the post points out, there is no way this is “the best they could do”. It’s a write up that didn’t involve someone with the math + communication skills required to clearly explain the result.
Nice idea.
(See https://agmai.org/general-sep29/ for the recommendation in question.)
People seem to have very misguided ideas about why OpenAI is doing this at all. It is not to brag or to torture mathematicians. It is an eval. OpenAI is known to be willing to pay large amount of money to get a good eval, think FrontierMath. FrontierMath is now saturated, so they need a replacement eval for math. Open math problems are actually a fairly good eval, although a proper eval is better (eg FrontierMath has known difficulty and have tiers from 1 to 4).
Mathematicians would prefer if OpenAI didn't use open math problems as an eval, but OpenAI is not obliged. I actually think OpenAI wouldn't point AI to open math problems if unsaturated FrontierMath Super Duper is available, as it just angers mathematicians, but such eval is not in fact available. Given OpenAI used open math problems as an eval, they could just throw out the result (this is in fact better as an eval since it will keep problems useful longer), but mathematicians preferred to see the result. So OpenAI released them.
This is a misleading characterisation of the mathematicians' position.
The very first paragraph of the AGMAI recommendations explicitly states:
"we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models." You appear to have acknowledged this by saying “Mathematicians would prefer if OpenAI didn't use open math problems as an eval…”.
The mathematicians did not ask OpenAI to produce these results. They explicitly asked AI labs to stop producing them in this manner. Their subsequent recommendations concern what labs should do if they have already produced significant results, not an endorsement of the practice.
Furthermore, the recommendation was not simply to release the results, but to responsibly release already existing results. Section 2.B, Step I, explicitly recommends "...labs that have AI mathematical output that is not understood by the people who prompted the AI systems", to search the literature for relevant prior work, provide appropriate attribution, and improve the exposition of AI-generated proofs before releasing them, rather than leaving this work to mathematicians afterwards.
OpenAI published the results on GitHub while still exploring repositories that meet the committee's guidelines. So they followed some of the recommendations, but not all of them and hence, did not release the results as requested by the mathematicians.
I do not think it is a settled matter whether this was done out of goodwill. This is because releasing these results as they were can benefit OpenAI more than releasing them according to the AGMAI recommendations. AGMAI recommended in section 2.B, Step 1.5 that "Each time a solution to a problem is released, it should be clearly documented how exactly AI came to be used on that particular problem. If many results are released at once, then in addition to the results themselves a further document should be written and made public that references all of the released results and explains how many other problems of comparable difficulty the models tried and failed to solve, as well as how the problems were chosen." If the results are released, it is easy to expect that the media will discuss the capabilities of the AI used in the work, as indeed happened. If this AGMAI recommendation was followed, the media would plausibly have also discussed the number of failed attempts and then the overall attitude would not be as favourable to OpenAI as it is now when it comes to the capabilities of the AI that was used. OpenAI did release on GitHub that approximately 4,000 problems were attempted and resulted in 719 manuscripts (after 3 containing suspected errors were removed by OpenAI) across 372 families of problems, but this does not give a calculable number of problems it failed to solve. I do not claim to know OpenAI's intentions or reasoning when these results were released and am not arguing that it was done with improper intentions, only that whether it was done out of goodwill is not a settled matter.
AGMAI's October 6 statement explicitly clarified that its advisory role should not be interpreted as an endorsement of OpenAI's process, and that it was up to the mathematical community to assess how successfully its recommendations had been followed.
Recommending how to responsibly handle the outcomes of something you oppose is not the same as asking for it to happen.
> Mathematicians did not ask for this work to be done. The Advisory Group on Mathematics and Artificial Intelligence, from whom OpenAI has claimed to derive its legitimacy, opened their initial advisory statement by saying that frontier AI corporations should not test advanced mathematical problems on internal models. In ignoring the central premise of the Advisory Group’s position, OpenAI has indicated total disregard for the norms of scientific research — norms that guarantee that mathematics remains trustworthy, ethically researched, and in the public interest.
That's right honourable of you but for me it is very clear that the only incentive in AI companies' effort to produce mathematical results is to advertise their technology. There is no reason at all to assume they have any other motive; certainly not any kind of interest in mathematics as such.
> supported by a clear plurality of respondents
was referring to?
People who are invested in the idea that we've invented a general intelligence, now, which includes all these companies that are literally financially invested in this claim they are making, will tend to believe that its results can already be trusted in domains like this. Some mathematicians seem to believe some of the proofs written by their models, and some, like this one, don't. I do think it's valid for an expert to push back against the claim that the best use of their time right now is to verify the poorly written work of everyone who's claimed to solve the problem
That's not true. [0]
> On July 25, Ramana Kumar published a repository containing a sorry-free "disproof" of the Collatz conjecture, produced with AI assistance. It is not a valid proof because it exploits a bug in the kernel's handling of nested inductive types.
Even in this dump we're talking about, it hasn't been true. [1]
> In “Algebraicity of Weil classes on split abelian eightfolds” a sign error invalidates a stabilization-trace cancellation argument and the construction used by two dependent papers.
[0] https://leodemoura.github.io/blog/2026-8-24-postmortem-for-t...
2. None of the results Open AI retracted had an attached lean proof
If you just strip mine the answers and Sam Altmans magic button solves 100/100 problems, what's next? Who is left to come up with a new interesting question for the magic button to solve?
Lastly, life and the present moment is all there is, if there is no enjoyment in anything we do, then what's the point of all the "living for ever" Altman et al want to achieve.
We will live forever to read boring papers generated by LLMs? Literally sounds like an eternal hell.
In fact, I can't remember a time when I was more excited about the future of science. This could herald an end to the replication crisis, and kill off bullshit science completely. The danger of course is that we end up with two companies effectively dominating cutting edge research in every field, but it remains to be seen if that's even possible given the pace of improvement in open weight models.
He's tracking the community progress on sub-n log n multiplication. OpenAI started with 1 - 1.63e-55. The result has been now improved on 115 times, and the current record is "rohanarun"'s 1 - 9.87e-5. I'm sure by tomorrow it'll have improved again.
Does this look like people aren't having fun? Does it look like they aren't discovering stuff? It looks like it's spurred a cascade of interesting community activity. It doesn't really seem much different from what happened with the twin primes conjecture. Isn't that supposed to be the point of all this?
This is like "no one is forcing software engineers to use AI tooling" or "no one is forcing you to show your ID in the airport" or "no one is forcing you to own a car in your small midwestern city" - there can be no law requiring something and the practical consequences of not doing so can be so painful that you're effectively forced anyway.
That’s why it doesn’t make sense to present AI companies as dumping or burdening the scientific community into doing labor for them; the scientific community is self motivated to do so.
As someone who uses LLM tech occasionally, this is why I prefer using open local models. If I’m making myself obsolete, at least I’m not making some asshole richer and their closed model better.
I just imagined that instead of math papers, they released 700+ feature length films, and the only way to tell if one of them is any good is to watch it in its entirety.
That feels pretty unappealing to me.
I know it's the same for human made films, so what's the difference right? But those are good enough most of the time that it's a decent bet, and the people that made them had real skin in the game.
Contrast that with something made by a nondeterministic slop machine with no skin in the game where small details can be off in a way that's jarring. Right out the gate I have an aversion to committing that much time to something that very well may waste it.
However, there are always smarter, hungrier people out there and this is a buffet.
Some output is going to be wrong or incomplete. I am willing to bet even those have nuggets that can be used elsewhere.
Like people enjoy racing in front of a stopped train? As soon as they turn on the engine again, they will run you over. The questions that remain will be only the low value ones, not worth the effort to vacuum up.
So no, the smarter, hungrier people are not the ones that are going to swoop in. It will be the most desperate.
> Some output is going to be wrong or incomplete
This is a very human take on the situation. No, the Lean proof is not going to be wrong, and it will be incomplete only in the sense that OpenAI didn’t try to push the results further.
There's a "Silicon Valley-ism" for you. We offer a thing in whatever form we want and people "who are passionate" will gobble it up, should gobble it up, 'cause they're "passionate".