OpenAI withdraws three mathematical results
twitter.com
twitter.com
What about software developer community? AI has eliminated the need for junior software engineers. Almost no one is hiring junior software engineers. But companies still need senior software engineers. Without junior engineers how will there be senior software engineers in the future?
What is the solution? I don't think the solution is to say AI progress in software, mathematics etc. should be halted.
Your code used to be a masterpiece, so well crafted it's easy for AI to tweak and modify because you've got everything so logically organised and scoped... and now this is what you're producing? Hard to debug monstrosities that only an LLM can realistically bolt new features or tweaks onto, because it can do the kinds of refactoring necessary each time.
We use to talk about the fact that code should be readable because you spend more time reading it than writing it, but I think that misses the key part that readable code is also typically easier to debug. If you can read and understand the code, you can follow the logic when things are wrong in production, and you can more easily reason about the emergent properties of interactions between the complex systems that are involved.
I jest, while in agreement with this whole comment. We used to care about fostering informed developers and maintaining high standards and good quality software.
I literally compared it to being an expert woodworker. Beautiful ornate decoration. Rich, sturdy mahogany, one of a kind, beveled edges and a fantastic stained hardwood.
Now it’s the 30$ Ikea cardboard stuff.
My biggest question is how many tables does the world need, and how many woodworkers will be required to build + maintain those table factories.
there are several features over the last few months that were obviously made and deployed and no one even launched the dev server and tried a single thing to verify if it was right. just "pull my ticket, do my ticket, push my ticket. i am a developer."
I’ve seen monthly and quarterly executing results AI generated with plainly wrong factual information.
IME a lot of people are now subject to output-rate expectations that preclude doing much else, honestly.
Ahh that's OK then. Everyone's in this same boat simultaneously in multiple industries! Cool!
> What is the solution? I don't think the solution is to say AI progress in software, mathematics etc. should be halted.
I think the solution from the maths world is to not grant these AI papers (or their human sponsors) the normal courtesies of "regular order", just as you would not with an AI lawyer or someone who was just pressing enter at a law firm.
But in the software world, nobody gives a shit, apparently. We are collectively morally bankrupt and should not be granted the regular order to help other people to decide what to do with us.
In fact, the world is always filled with curious people who like to go one level below.
This hysteria about losing "Junior Software Engineers" -- most of them in it for money, promotion rather than craftmanship, is over-rated.
People who love solving puzzles will always find ways to sharpen their mind.
People who love understanding things, will always find ways (AI will help them tremendously).
People who love taking shortcuts will always find ways for it (AI or not)
If you have made it to the point of being a Junior Developer, I can assure you food and shelter is not a problem for them. You just have to adjust to a standard of living like the other 7 Billion people on this world.
Also, if a Junior Developer can show me(or anyone) they have built an entire system on their own and explain key concepts, there is no dearth of jobs for them
Eh? Apart from it not being what they are paid to do on their 9/9/6 jobs, when will they have the time to make it happen?
What is going to happen is that the remnants of the open source community will do the job of educating juniors for free, when the university degree system collapses. Just like it currently keeps a bunch of systems going with inadequate compensation.
The corporate world gets the problem off its balance sheet. Again.
The actual argument is "there will be no jobs for Junior Engineers, so there will be far fewer, and as a result there will be a huge shortage of Senior Software Engineers".
You are responding to the problem as if it some kind of extinction event, like a rare bird, where if we can find a breeding population we save the day. A few curious people, self training for the love of the game, and as a result we still have a few Software Engineers so everything is fine. It is not like that, and I haven't heard anyone suggest that is the issue. The potential problem is a massive shortage of workers with skills that are currently essential to the functioning of a large fraction of the economy, whom we might still need in the future.
The continued existence of talented enthusiasts does not establish an adequate workforce pipeline. If paid entry level experience contracts, what replaces it, and why should we expect that replacement to operate at sufficient scale?
You have to do the hard thing eventually, or you never get anywhere.
AI progress in software
Oh, yes... the progress... You measure it by LOC, right?Some jobs will stick around in vastly diminished numbers with tasks that are completely different than what they used to be to produce the same output (e.g. farmer). Other jobs will be eliminated entirely (e.g. switchboard operator). I'm guessing things like software engineering will go the way of the farmer, with the main unknown being just how much demand for software there is.
The entire valuation of the AI industry is predicated on people not just losing individual jobs, but being taken out of the workforce entirely on an economic level.
They are talking about the workforce of the entire economy shrinking. People will lose their livelihoods for good.
This has the potential to be even worse than the second agricultural revolution to industrial revolution phase, which made ordinary workers lives absolutely miserable for maybe a hundred and fifty years.
This time, there will be no jobs. If you are displaced from one industry by AI, you will end up in another industry also being decimated by AI; if you get a job at all, you will do so by working lower pay than other workers, who will in turn be pushed down the ladder.
And that is if you are lucky: if you have only IT skills, why should you be the first to get a fruit picking or plumbing job?
In this case with AI that power will shift to the companies that run the AIs
My take is: the bar to what counts to HR as "senior" will go down as businesses everywhere try to adapt and hire more seniors - "senior" now just a name, as it becomes the new "junior". Then everyone will pat each other on the back till it all goes down in the flames of bankruptcy.
Some people hope AI will get good enough in a few years that it can innovate without human experts. Maybe? But that remains to be seen.
By the progress of AI from ChatGPT to now is horrifyingly fast.
Trendlines and basic reasoning point to a most probable future where the AI is superior. We can’t just say “that remains to be seen” because the alternative is the least probable future.
Anticipate the change and act prior.
Right, companies won't need software engineers. They'll just need someone who can use tools to produce source code and maintain the generated artifacts, plus make domain-specific technical decisions like "what should the system do when two users update the same record as the same time" or "how should the system behave when a message in the queue cannot be processed".
We really oughta come up with a job title for these people.
agent herder!
If that comes to pass, I will have to re-evaluate my career options
I do not know or care what my if statement turned into in x86 assembly unless it becomes a performance problem and even then, I'm not profiling or debugging in machine language. Neither do most developers these days. A message in a queue becomes something akin to that in this era.
I see that I got downvoted there. This is not something I advocate or look forward to but I feel this is where it is going.
“ Claude! what should the system do when two users update the same record as the same time, explain to me with full clarity”
or "Claude! how should the system behave when a message in the queue cannot be processed? Give me all the possible ways ranked from best to worst, also explain to me all these concepts so I can understand as I don’t have a cs degree".
If you think that there is no future where software engineers don’t matter then you are delusional. While the future is not set in stone the pace and trendline of AI point to this future as a MORE realistic future then the alternative.
Your example btw is ALREADY a solved problem. AI can answer it and design around it. Agents at my company already handle our infra.
I have a data pipeline with 6 steps, A -> B -> C -> D -> E -> F. I asked Codex to make some specific optimizations to step B and benchmark them. It did what I asked. Then it decided to also benchmark the entire pipeline, and after noticing that step E was slow it decided to make some optimizations that I had not asked for on step E. It was at this point that I wondered why it was taking so long, saw what it was doing, and stopped it.
This is GPT-6.1 Sol High.
I think that person does not need to know about locks anymore.
In the age of extraction capitalism where building sustainable, profitable companies is not the goal, no one will care.
In other words, they are increasingly devauing their own senior position and discarding their hard earned skills that make them seniors in the first place.
There will be no more seniors. Tech companies will hire junior llm agent wranglers who took a class in undergrad doing this. That is probably the nearterm.
And that's a bad thing. If the math community didn't exist or was weak, OpenAI would still benefit from the prestige of these results they were forced to withdraw. Withdrawing these papers has harmed OpenAI's investors, and that's totally unacceptable.
> With automated math that community as tao pointed out is at risk.
Good to hear. The problem they represent needs to be eliminated.
This concerns the Hodge conjecture (millennium prize related) paper. Seems to me like PhD nerds weren't confident bosses pushed ahead anyway.
1. This is expected if you only use a single model family like Claude, eg. we use a different model family for code review than authoring, OAI could have done this too for their math dump
2. Ai needs a good human driver beyond the trivial or mundane, they are expert enhancing machines, not expert creating machines. This is where the community comes in. Reading Tao's ChatGPT session reveals this: https://news.ycombinator.com/item?id=49010345
3. OAI is not trying to be a member of the/any community, this is not the first story to shows this, nor do I expect it to be the last. Perhaps this is them being effective altruists today? /s
I agree LLM review is also fallible (as is human review) but the interesting part to me is that finding this sign error before publication should have been table stakes for OpenAI, it’s their own model that found the sign error.
I’m curious what was in the original prompt and what was in the prompt that led to finding the sign error, I think it matters a lot for understanding the dynamics here
As much as anything can be infallible.
If they can be automated, they are not necessary. If they are necessary, they won't be fully automated. It's a pretty simple experiment to run, the math "community" should bear with us. Darwin would be proud.
Even if your stated assumption was baked into the original comment, which is doubtful: the historical record shows that we will keep relearning The Bitter Lesson and each community will pretend what they do for a living is exceptional and immune because of xyz. The screams will get louder when the "greedy" and "dumb" automation comes knocking and it turns out nothing was truly immune or "nuanced ".
Getting some new hobbies may be in order, it's a Brave New World.
It's great that we're starting to see the light at the end of the tunnel, and will some day achieve a perfect market without humans. If you think about it, all the market really needs is a people to own everything, everything else can be automated, and all those annoying human workers can be eliminated.
To witness an arson and rejoice reveals an ugly kind of sadism.
It is pointless to call out the little wins humans still have because again, those wins are temporary.
We need real concerted effort into asking: what is the point? For me the only answer I came up with is: fun.
An absolutely ridiculous statement. There is a vast amount of mathematical knowledge that hasn’t even been written down, much less formalized.
You’re distinguishing “knowledge” from “idea” in a particular way that doesn’t correspond to common usage (see my counter examples). Without you being explicit about your definitions, I can’t tell whether what you’re saying is meaningful. It feels tautological.
Given that an executive assistant has unwritten knowledge that is necessary to do their job, where your evidence that no mathematician has analogous knowledge (using the word in the common way, not whatever way you mean it)?
It’s possible, but it’s not as obvious as you seem to think.
If so, why did they mix proofs that were verified with Lean, and proofs in natural language?
I was wondering that while reading Aaronson's blog:
https://scottaaronson.blog/?p=10169
Or at least, we’re pretty sure that it’s a proof! There’s a Lean certificate, as there are for some of the other 372 breakthrough results (not all of them). But it also appears that no human has understood just about any of these proofs yet
It seems that the obvious thing to do would be to release in TWO parts: the ones that are verified, and the ones that might have some good ideas but also might have some mistakes. Presumably the latter would be much more epxensive for humans to verify.
Can you explain this? How would having a lean proof of the program make it more likely the proof is weak?
https://leodemoura.github.io/blog/2026-8-1-postmortem-for-ke...
There's also a fairly well known incident where the Lean formalization of the Riemann Hypothesis in Mathlib was incorrect.
I'm probably wrong, but what's the point of throwing away all skepticism?
> All the human needs to verify is that the statement of the theorem is translated correctly from natural language to Lean. That usually covers a very small surface of the Lean code.
That's exactly what the navier stokes paper posted yesterday pointed out where the LLM bends the Lean code to make it "compile", because the NL might be wrong to begin with or because it missed a detail:https://arxiv.org/html/2610.08144v1#S2
For complex / tedious proofs I can easily see how small details like this can lead to a valid lean proof (or valid "code"), but missing the important details that got lost.
The proof explicitly hand-waves some complexity by assuming lookup tables to avoid some calculations which isn’t actually possible since it’s dealing with such large numbers and it only works on incredibly large numbers.
The complexity being so close to nlogn and the handwaving by assuming lookup tables in parts should be a really really obvious smell. At the very least worthy of holding back from the broader announcement.
It us proven in lean as-is with these assumptions and it’s not one of the ones retracted but those assumptions are doing some heavy lifting. I think it’s worth adding back in those ‘by using a lookup tables for x’ complexities and seeing if we really are below nlogn on that one.
This is just a Rice Theorem problem, right?
2. Why shouldn’t math progress happen in the open, commit by commit? Why is it so horrible if a proof is 95% of the way there but we later find that it needs to be refined? Mathematics previously was optimizing for an antiquated publishing and distribution scheme. There is no need for the first print to be correct. We have the internet now. We can and should publish incomplete results and correct things on the fly. Maybe mathematicians would have solved some of these problems years ago if they didn’t hide incomplete almost solutions in their filing cabinet because it wasn’t yet ready to be published.
You don’t hate the pageantry of mathematics and academics enough.
I bet you enjoy when a peer asks you to find the issues in a fully AI generated PR that’s 95% of the way there.
Also “real mathematicians” aren’t the people who “math belongs to”, you’re a mathematician if you do math, that’s it.
There's a lot to hate about the academic world, but the solution isn't spewing out terabytes of crappy half-baked results.
The lean proof uses these assume ‘a lookup table’ assumptions. The paper smells with the nlogn^0.99999999 (many more nines actually) and unbelievably close to nlogn statement and then the literal talk of lookup tables pushes it over the edge clearly for me.
Maths can generate weird numbers out of nowhere but it really really looks like an nlogn result with some tricks to get past leen to me
Of course, that's not to say the research is necessarily useless. It's still theoretically interesting to find "better" algorithms if only to shed some light on lower bounds, and so on. And who knows, maybe the line of research could lead to more practical algorithms later on.
From the "Introduction" section of that paper: "The constants and thresholds in the construction are extremely large".
(And verifying if the algorithm multiplies correctly or not is the less-interesting part of this, anyway. Gets you no closer to verifying the complexity result).
Regarding elegance, take a look at Graham's number. It was not some meaningful constant - it's just a big-ass number which could be used in existence proof. Human mathematicians have been using this approach for quite some time, it's not really AI doing things odd
With 40% formalized they probably have a good idea of how many were found to have fatal issues in the formalization attempt, and they hired some mathematicians to verify some of them, especially the big headline ones.
Are they all too busy having brilliant ideas? Doubt.
OpenAI math paper dump should be considered like a hint from 200 IQ eccentric genius - unreliable but perhaps insightful. If it was not "ugh AI" people would be happy about it.
It would be ridiculous for anyone to say "hey man you really shouldn't even have posted these unless you have an ironclad proof".
I'm curious to know if the withdrawal was due to an actual mathematician looking at the papers and noticing the errors, or they ran a model on these to proofread, which would not be the first time, presumably, since they would have surely done that before publishing. Both options have interesting implications.
The latest model even finds mistakes in previously published math papers!!!*
* so far, only OpenAI's math papers were faulty and needed retraction.
It'll be a good test to separate those earnestly trying to advance human knowledge, from those wasting my tax dollars. The later group ought to be publicly shamed and ridiculed without mercy. We need a more invective word than 'pseudo-intellectual.'
I don't have a problem with them publishing. I don't have a problem with the process and how they are interacting with it. I am delighted that they are actually acting as stewards of these works.
All that aside, they should be paying the people verifying the problems. The thing that really gets me is that we know anything published in the process of verifying this is going to be vacuumed up into the next training session.
It kind of reminds me of when tech giants open source a project as a means of putting a positive spin on abandonware. “Here’s the source! Any problems are yours to fix now. You’re welcome”
I also fail to see the issue you have with releasing abandoned source. In what world is that bad? That obviously is a gift and should be encouraged. e.g. id software's history of doing that has meant their work stays alive forever.
Ironically though, what I imagine will happen, is that the researchers will pay OAI to use chatGPT to help themselves eval the proofs.
This _might_ have been true somewhat in the past (although it wasn't), but it's completely false today. Anyone with access to a sufficiently advanced model has the capabilities of analyzing these papers/proofs. It's no different than reading a codebase you might not be fully familiar with, and checking it for correctness (give an engineering analogy).
This hardcore gatekeeping of math (and by extension STEM) fields MUST stop.
Like I was reading some about adele rings last night, which is already going to be quite a concept for a layman to be able to even slightly describe. Then you can layer on that apparently they're locally compact, so we can talk about harmonic analysis on the additive group. Like, come on now, 99.99% of people have no hope of ever following along, and this is stuff from 75 years ago.
But they don’t. What they have is the ability to ask something else to do the analysis. It’s an important distinction. If the asker has the skills to evaluate the results, that’s one thing, but too many don’t and act as if whatever they got is unambiguous truth.
> This hardcore gatekeeping of math (and by extension STEM) fields MUST stop.
What must stop is the overuse of the word “gatekeeping”. Anyone is free to study these fields and work on problems. What people rightfully object to is uninformed research flooding everything with hard to verify junk.
Withdrawal is akin to submitting a paper to peer review and then when you’ve noticed mistakes, you decide to take the paper back and correct it.
Reject is when someone else notices the mistakes and tells you to take it back and correct it.
Withdrawal and reject happen all the time in a scientist’s career. They don’t necessarily mean the scientist is doing bad research, just the research was not ready. Retract usually means something more.
By dumping the papers, OpenAI skipped the typical peer review process, so peer review should be understood as what’s going on now as mathematicians look over the papers and find flaws.
IMO If you take out all the stupid human aspects mostly related to fear, egos, etc, we should brace the imperfect and helpful tools, whatever they are, improve them so they are as easy as possible to review, and keep that core scientific discovery loop going
This can and should erode our trust in every single proof OpenAI published. The model is clearly faliable despite the lean proof, and clearly the output was't actually checked properly before release. Once these proofs are peer reviewed and published in a journal we might be able to trust them again but until then they are just slop, sadly.
I think OpenAI actually did the right thing by sharing everything with the whole community right now but I also hope that some significant credit will now go to the reviewers who confirm these 'proofs" actually work.
Humans produce flawed Lean proofs. Indeed LLMs were successful at finding and fixing many issues in the “core” standard library if I recall correctly. Humans regularly produce flawed papers and have minor issues require fixing. And when it happens it often isn’t as prompt and clear as this.
In fact we already follow exactly this process for human papers and have done for a very long time. Publications without peer review are treated with great suspicion. This is how we end up with journals of varying levels of prestige and rigour.
The process isn't flawless and there are huge problems with retractions, as well as weird financial incentives and rent extraction but there is definitely an increased level of trust in a paper published in Nature.
More seriously the problem is the complete utter lack of care OpenAI has shown in their desperation to demoralize mathematicians with their new LLM. In their words, it took 3 hours of ChatGPT pro per result, why not spend a hundred hours per result formalizing it, checking if the formalization matches the natural language proof, and whether the argument could be made more simpler and readable. Any human paper has hundreds of hours of work put into it, but OpenAI who is absolutely adamant in demoralizing the mathematical community and demonstrating their superior “intelligence” will only spend 3 hours, write unreadable, inscrutable proofs, not formalize all of them, and then dump it on the mathematical community for some reason.
But how much more progress would have been made in the last 100 years if they collaborated in real time? Watching the work someone is doing and spotting errors, or making suggestions would be a good thing. Unless the goal is simply to claim credit for a discovery vs the discovery itself.
If someone had a blog about their research on a math problem, and they figure out one day that what they wrote a week before was wrong (retracting it), would that be a bad thing? I don't think it would be bad, but that's not how academics approach things.
> As part of our GitHub repository, we are sharing formalizations of many of the proofs in Lean, a programming language that allows mathematical proofs to be checked by a computer. We will update the repository with more formalizations as we obtain them.
Meaning they published all results before checking all of them, and intended to add more Lean proofs later. In the linked post they state ~42% of the posted results now have formalized proofs, some were added, some verified, and I assume this means that some results turned out to be wrong.
If your AI tool can help advance mathematical research, share the tool with mathematicians. Using it like this is irresponsible.
"AI will kill us all": no. Greedy humans will kill us all. With AI.
You might think this is not very useful, maybe - but that’s not a reason to retract..?
If you want to see paper retractions, you ain’t seen nothing yet.
That’s almost certainly due to things like captcha or something. What’s the actual shop you are referring to?
This heuristic indicates that at most a handful of them will be of marginal value.
> The vast majority of results were obtained with the same procedure using an unreleased internal OpenAI model. On average, each result used three hours of ChatGPT Pro thinking compute with that model. Over the course of the evaluation, the model was posed approximately 4,000 problems. Aggregating the output into result families and manuscripts and requiring an appropriate level of significance led to the catalog outlined above.
seeing the full list of problems would be the most interesting part of this whole situation. it could give some insights into what kind of attributes of problems cause issues / are easy to solve for LLMs. (edit: they posted results for ~700 of the 4000)
This is a "complaint" that Tao had (I think it was on his blog) is that if mathematicians could see the failures, it might provide insight of where/how the models struggle. Of course, it's unclear if these failures can be addressed with more chips/training/etc.
Is this of practical use, or just a proof for now?
Look at examples here: https://en.wikipedia.org/wiki/Galactic_algorithm
If that’s the case in a way its a similar delusion that average people are experiencing with their own AI use.
You could almost draw a comparison between that and inexperienced consultants making business changes, claiming glory and then disappearing before the thing falls apart.
There is however a new problem of scale. Erdös was a human and still managed to create work for an entire generation of mathematicians, how much of a mess will an automathician create?
I vote for "automathon"
Interesting that math gets so much attention, when actual advances to material science, biology and chemistry have much higher ramifications and economic benefits. I assume progress there is kept under wraps until they can capture the economic benefits. If they can do that, then the insane valuations may actually be valid.
Or progress is not as straight forward in those fields as in math.
> The repo now has ~42% top-line results formalized.
Understanding the proofs is a different story unfortunately.
We let the experts investigate. If the results are dodgy, then the next batch of results will have to do more upfront work to demonstrate their worth. If there is gold in them hills, then this is exciting though very disruptive for the math community.
* The reason you shouldn't consider the withdrawals to be caused by errors is because this is pretty standard in math and development. "Errors" like this are a core aspect of science and it happens _all the time_. And from my research LLMs have a far lower error rate than even the best human scientists.
Even Einstein retracted one of his earlier papers re the cosmological constant... And he is one of the greats. Although not purely math related, it still counts.
Or is it the kind of research you'd prefer to keep shrouded in mystery?
Or if the "peer review" holds up for the remaining results, then it's fair to say that the AI hype is real and the world is about to change dramatically and faster than anyone can comprehend.
So which is it?? LLMs can do some really impressive coding. Bug fixing. Exploit finding. It has reasoning abilites that advance every day. Solving real math problems like this is one thing I was waiting on. It will be interesting to see if it holds up.
If it does, we should expect many other advancements to follow in many other areas. Disease, material science, fusion?
I mean, even if just a few results ultimately hold up to scrutiny, isn't that something that would have been regarded as a major advancement regardless of if it was AI?
The cynical view still makes me think that at the end of the day all the models can do is predict the next word. And as a result, they will be very limited to certain tasks like coding. Math reasoning is much different from writing code. Time will tell.
Can't be, as IPO isn't until next year, and I'm pretty confident that is there are errors, said actual mathematicians will find at least a few in the remaining ~2.75 months left in the year+. If there really are errors and a decent amount aren't discovered before the IPO then I'll have serious questions about the math community.
there is also will be new option of "maybe correct LLM proof", which humans will never be able to comprehend and verify.
AI is this strange modernist machinary that kind of threatens that brhaminic role... its almost like the vatican vs post industrialization world .. where they still have to keep making the case for why religion/priesthood/god is important... even as the tech/science world starts operating on totally different terms...
This metaphor might be applicable if the AI slop machine was in fact producing novel output. It seems to be getting invalidated as people dig through the wall of meaningless text surrounding the actual results.
3 results out of 400 is nothing.
It just hallucinates an answer and then makes up workings to go with it! Just like when they start hacking and lying because the problem is impossible…
No it’s not. In fact, much of it is formally verified, which makes it far more reliable than most human-written proofs.
Btw, the most famous human-written proof of the past half-century (Fermat’s Last Theorem) had a massive flaw that took two years and major help from other mathematicians to fix, while the most (in)famous human-written proof of the past 15 years (abc conjecture) is now widely believed to be false.
But people hear what they want to hear I guess.
its formal verification on top of formalization by LLM, which could have errors.