For example, had they decided to put their paper on arXiv, they would be facing a 1 year ban, followed by not being able to put anything on there without prior peer review.
965 karma · joined July 31, 2015
For example, had they decided to put their paper on arXiv, they would be facing a 1 year ban, followed by not being able to put anything on there without prior peer review.
> During this event, OpenAI requested an interview concerning my vision of the future of AI and mathematics. I accepted, and spoke with them for perhaps an hour. I had done similar interviews in various venues, and I assumed that, as with these other cases, they would eventually post the entire interview online, which talked about both the possibilities and risks of AI much as I have done in these other interviews. As it turned out, they only used a few snippets of that interview for that infamous advertisement instead.
https://terrytao.wordpress.com/2026/09/07/finite-time-blowup...
Similarly, if all you ever read were OpenAI blog posts, you would get a very wrong impression of the usefulness of large language models of today in maths. For a working researcher, it's not a magic wand that you point at any given proposition and it tells you whether that proposition is true or not. It does appear to help if, while pointing your wand and utter the magical incantation “do it up bro”, you also make it convert $15 million into heat, but for most people, this kind of inverted Midas touch isn't quite accessible yet.
Instead, the reality seems to be closer to this, projecting a fair bit: a given mathematician will have a collection of propositions that they care about, and that they'll use as their own internal benchmark as new models come out. Very rarely will anything come out of it, but sometimes, in particular if you make sure to provide the wand with all relevant context, papers that could be relevant, proof strategies and lemma structures that you suspect are useful, something (which may or may not be plagiarism) will pop out, and that's really nifty. Moreover, it is not unimportant what the proposition and the relevant proof is like. And what does come out tends to be quite bizarre; proofs that use terminology that doesn't exist, seem overly pretentious, based on nonsense analogies where it's surprising that it even works at all, and the only comfort is that you can join it with an equally unreadable Lean blob. And where you would be _crazy_ to just publish those artifacts and think that you have contributed much of anything to maths.
But sometimes it works. It's still very unclear what kind of maths the models are good at, but it seems to certainly be an advantage if what you're looking for is a counterexample hidden in a pile of otherwise similar-looking non-counterexamples, if your proof is one that requires considering 36 different cases, each of which are so tedious that no researcher would have the patience to go through them by hand, or if the proof is an amalgamation of several existing structures, some of which are only documented in Georgian.
The gold rush, more than anything else, seems to be populating the convex hull of existing maths.
This can all change. The $15 million wand requirement today will be less tomorrow. Whether we ever get a move 37 is less clear, or whether we will eventually reach stagnation as all low-hanging fruit is picked, and the convex hull is populated; call this cope if you like. But maybe we do get move 37s all over the place, and it's fine that people think about what that future will look like.
Until then, and while we're still picking friut, let us rather have a think about what we can do to fix the incentive mismatch, to ensure that we increase the prestige of digestion over being the first to convince the LLM to do it up. Since that's the one thing everyone seems to agree, chances are it'll probably converge to something that doesn't have to be written in commandment form, but out of the guest posts hosted by Tao so far, the one by Antieau has some useful suggestions for standards (that aren't entirely unlike those from Leiden): https://terrytao.wordpress.com/2026/09/15/fast-math-slow-mat...
Part of the point of the letter is that it is quite possible to act in a way that is a net negative to research. The most obvious case is when the companies violate ethical standards in research.
The subtler case, the one for maths in particular, is what happens when you fail to follow well-established patterns for making maths research productive. Tao himself spelled out how that can look in https://mathstodon.xyz/@tao/117207856734787448 (which notably came before any of the news on Navier–Stokes).
So let's just quickly agree that the actual quote is “However, the push by AI companies to solve mathematical problems as a benchmark is detrimental to the science of mathematics, and to the mathematical community.” And that in this, "as a benchmark" is load-bearing.
We wouldn't see nearly the same amount of contempt from researchers had OpenAI picked a research-friendly approach.
What they did: Hear a rumour about the problem being solved by other researchers, then rush to scoop them (unethical), then, when they actually go talk to them, they try to oust an author (also unethical), and when they finally decide to share their own work, do so in the least useful way possible.
What they could have done: Upon hearing the rumours, connect with the researcher in question and propose that they join efforts instead; set up a joint project to test if the machines are useful in any way, and if that's not appreciated, back down again. And instead of dumping only an undigested paper* and a Lean proof, do the digestion prior to publishing anything (as Buckmaster was in the process of doing). If their own lack of competences was keeping them from digesting it, then again, reach out to the researchers to understand if anyone would be willing to do so.
In the second of those two worlds, we wouldn't be seeing nearly the amount of outrage that we are seeing right now.
*: Here, “digestion” is the process of turning an AI slop paper into something humans can read. LLMs can indeed sometimes (if much more rarely than marketing material from the large LLM companies will suggest) produce correct proofs, but they are often written in bizarre ways – they'll use lingo that doesn't exist, seem overly pretentious, dwell on extremely easy steps while glossing over the hard ones. Currently, a real researcher will take that output and transform it into something that others can understand, use, and build upon. This is not so different from what happens when using it to write software, although as someone who does both, I will say that the amount of digestion needed for proofs tends to be orders of magnitudes larger than for code. This meme is quite accurate: https://mathstodon.xyz/@tao/117068266071803252
This has a few practical implications: First of all, if you are in the target group of the marketing material, be wary. While these things can do non-trivial stuff, the amount of magic is being grossly over-stated. But also, when several of the big results have indeed been reappropriating the work of others; when the companies fail to provide proper attribution (the NS case in particular is laughable) and present the results as the models' own work, that's plagiarism.
I can see where that's coming from, but I really don't think it's the case. Even with Astra, the proofs you get are just off in a way that doesn't signal superhuman comprehension. As 9question1 says, a common theme is that they dwell on insignificant steps. Another one is that they'll often be full of terminology that either doesn't exist, or has this weird quality where it looks like it is trying to make some minor insight seem much greater than it is. At first glance, that'll often make it look like it knows more than you, but when it's really just doing the same thing but in a more complicated and worse fashion, that to me isn't a signal of comprehension at all. The bizarre thing is that despite all the "stochastic parrot" style nonsense you'll get in individual proof steps, they still often combine to something valid.
In either case, what all of this means is that the working mathematician still needs to go through, and generally completely rewrite, any proof output by an LLM. Otherwise you are passing the burden of unreadability onto the reader.
> Often these solutions are announced in a rush, leaving no time for a proper writeup, the isolation of new methods and ideas, and citing relevant previous work of others. As in all creative professions, this raises severe attribution and plagiarism questions.
Tao is not someone who is anti-AI for the sake of being anti-AI. He has been advocating for the usefulness of AI in maths for a long time, to the point that people have started calling him a shill for the commercial companies.
And everyone agrees that there are plenty of use cases to be had; helping with less interesting tasks like easing literature review, efficiently delving into existing work, doing review, whether on your own work or that of others, prototyping algorithms in areas where computation is useful, but also more in hands-on aspects of maths like validating potential proof directions by getting quick feedback on veracity of lemmas, etc., and, on very rare occasions, being able to one-shot the problem you care about.
The point he is trying to make here is much more subtle than "AI bad", and it's probably easy to miss if you have never engaged with research in maths: it's that the particular approach that large commercial companies have opted to take to produce marketing material can be a net negative. There is not doubt that -- even if you ignore the rampant plagiarism that has been reported across multiple problems now, the unethical attempts to oust authors, the outrageous attempts to scoop researchers instead of collaborating with them and building on existing projects -- it's nifty to have a machine that can help you figure out if a proposition is true or not. But just figuring out as much was never the point. When people have built problem lists, it's because some problems are more likely than others to provide new insight, and that insight is the target. And to than end, a poorly written paper with inadequate references and a pile of Lean is not valuable at all. Yes, now we know with higher certainty that Fermat's Last Theorem is true, but everyone expected that already.
One place where "just" answering the question can be a net negative is because the current incentive structure is set up in such a way that going in afterwards, trying to reclaim and extract the insights from a brute force solution, is considered less valuable work than that of coming up with a solution in the first place. That's a problem of incentives, and something Tao himself has addressed in e.g. his ICM talk, and that's something that we'll want to do something about. Until a better structure appears, though, if any given commercial provider of large language models really wants to help out with maths research and not just make more pre-IPO marketing material by competing with their customers, they could do so by using their magic machines to help build insight instead.
And with this paper, it looks like it may have come from an unpublished but publically available draft:
> Supplemental AI statement of the third author. Very unfortunately and without me being aware, the in-progress version of this work dated January 14, 2025 was shared publicly at the link https://www.ihes.fr/~gabber/BGV160.pdf on January 16, 2025. My understanding is that after 18 months it became a feed of the AI models. On July 23, a graduate student sent me the 37 pages pdf file at the link https://people.math.binghamton.edu/adrian/AI.pdf produced by AI when prompted on how the 7 pages AI paper https://www.ulam.ai/research/jacobian.pdf relates to BGV160.pdf; a typo in the statement of Theorem 30.1(2.j) was fixed based on what the file mentioned (its application to Theorem 30.1(10.c) did not require any change).
> To my best knowledge:
>(1) there has been no prior work before BGV160.pdf aiming to classify all étale endomorphisms of the affine spaces that have geometric degrees 3, see Section 26 of BGV160.pdf
> (2) there has been no prior work before BGV160.pdf that used systematically xy, 1+xy, and δ +xy in the construction of open embeddings, morphisms, and étale endomorphisms, including the morphism A2 K → P1 K defined by the rule (x : 1 + xy);
> (3) between BGV160.pdf and the 7 pages AI paper there exists overlap of notation (such as Γ for complements) and of intent (such as the variation of the geometric degrees, counting fibers, non-proper locus, families, the usage of 1 + xy, the computation of the normalization Xe which actually plays no role in terms of the counterexamples to the Jacobian Conjecture, etc.) which are also very pertinently and systematically presented in the 37 pages pdf file;
> (4) as of today the prompt history, the model version, and the workflow behind the 7 pages AI paper have not been disclosed publicly.
And why aim straight for scooping other researchers upon hearing rumours about their success? Normal, ethically acting, researchers would never do that.
And how about existence of non-sofic groups, which is actually the topic here?
While his concerns expressed here aren't new, it looks like the recent cases of scientific misconduct were enough to push for urgency:
> But in the immediate term, the most pressing issue is for the entire mathematical community to unite around our core values and objectives, and reject irresponsible and unsustainable usages of AI technology that only serve to advance nominal goals rather than the true underlying goals of the field.
Normally, when you tell a coworker that you're wrapping up a result, unless they're some kind of sociopath, their natural inclination would not be to try to steal it from you.
Your post says “While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .” We can discuss what it means to “read” things but obviously the issue here isn't whether you did it manually or automatically.
But more importantly, what on earth are you doing threatening real scientists to remove their coauthors, then making fun of them on social media? Does the entire company run on that toxic culture, or did those people run off of some kind of outrageous tangent?