Solving 20 Erdős Problems with 20 Codex Accounts Running in Parallel
starfleetmath.com
starfleetmath.com
Second, the proofs -- I understand the Lean 4 proofs to be refereed by Fable, and generated by Chat 5.6 Sol. Unlike the leaked proof of the Cycle Double Cover Conjecture last week which had a very nicely readable nearly humanlike writeup, the proof summaries (from Fable) read like Claude tends to read to me these days - real difficulty with the theory of mind of the reader, they are filled with technical phrases, acknowledgment of hard bits and oblique reference to solutions. In short, they suck. I didn't see the word load-bearing, but I bet it's there.
That said, a Lean 4 proof is a pretty compelling output artifact. I find it interesting that it's an additional type of effort to turn these into human readable / appreciable / beautiful / non-shitty proofs.
To those who say who cares -- indeed. But. One of the major reasons things like the Erdos problems are valuable is that they can at times spur new techniques and concepts. The best of these concepts are applied elsewhere, advancing the frontier. While we gain a lot from solving these problems, we'll gain even more from that next step of distillation / explanation into something humans and computers can grok together. I'd hope that with so many tentatively marked 'solved' we will see some new techniques / ontology / concepts. If not, still pretty amazing.
Are you running tool calls that include inference with local fine tunes? And fast math packages? Controlled by the frontier model agents?
Is there a way folks can contribute to this?
1) As far as the AI models go, we used GPT 5.6 Sol, Fable 5, and Gemini-2-embeddings across the system
2) Yes, the agents are given bash tools that allows them to interact with the preinstalled mathematics packages/dependencies that are on the VMs
3) This was a setup as a relatively quick project without much thought for future contributions, I will spend some time thinking about how i could make it more open.
I hate to bother with more questions, but I'm just so curious about this.
If you could roughly sketch out your agentic harness loop in a sentence or two, what does it look like? Which model(s) do the driving? How is progress measured?
What's your daily/monthly budget for this look like, if you don't mind my asking?
I also had this sort of thoughts when finishing my master's degree. I guess what breaks the cycle is that proofs (like other artefacts in other human activities) deliver aesthetic bliss.
But even then your argument breaks down when you consider more complicated cryptographic structures.
Well, programmers couldn't devise such a method either. You seem to assume that mathematicians are better than them at finding efficient algorithms, which isn't very plausible.
> But even then your argument breaks down when you consider more complicated cryptographic structures.
Even if help of mathematics is needed there, applied — rather than pure — math might still be sufficient.
And surely you don't mean all pure math, so your argument ends up being circular --- useless math is useless. You happen to think that this math is useless, but you might be, or turn out to be, wrong.
Also, you're misusing the term "self-referential". You seem to mean that it's a closed system ... but it's not, since mathematicians interact with it.
Finally: so what? We do all sorts of enjoyable activities with no benefit other than the enjoyment. Solving Erdős problems seems at least as justifiable as finding the trillionth digit of pi or writing an Apple ][ emulator in Brainfuck or playing those video games that you demonize by calling them addictive. (Compare to, say, compulsively reading Britannica's The Great Books of the Western World ... it's snooty judgments all the way down.)
Postscript to finally: This framework has wider application ... it's not limited to Erdős problems or anything else that someone happens to consider to be useless.
I don't think that's true for a reasonable interpretation of "often". I'm pretty sure the vast majority of pure math research is and remains useless.
> Also, you're misusing the term "self-referential". You seem to mean that it's a closed system ... but it's not, since mathematicians interact with it.
Then which better term do you propose? The point was that its subject matter lies within itself, which is very different from science and philosophy.
> so your argument ends up being circular --- useless math is useless.
It's not circular: You can replace "pure" with "useless" and the argument stays the same. I contrasted useless math with useless science and useless philosophy for a reason.
> Finally: so what?
I'll just point out that "so what" is not a counterargument. If you agree with my point but find it unimportant: that's fine with me.
> Solving Erdős problems seems at least as justifiable as finding the trillionth digit of pi or writing an Apple ][ emulator in Brainfuck or playing those video games that you demonize by calling them addictive.
The difference between achieving a new record for a video game, and pure math research, is that nobody is confused about what the former is: It's a game, or a sport. Speed runners or chess champions or pi digit calculators don't confuse themselves with noble researchers advancing the frontier of human knowledge.
> I'll just point out that "so what" is not a counterargument.
It really is.
> Speed runners or chess champions or pi digit calculators don't confuse themselves with noble researchers advancing the frontier of human knowledge.
Yow.
I still like doing maths by pen and paper, but this is fun too.
I designed and stood up a sovereign inference/compute on my intranet. It uses a trust model that allows for a controller (me) to spin up untrusted inference/forge machines for Lean, Sage, or other runtimes. Untrusted sandbox workers integrate directly into my custom harness as first class “attachments.” This is the “orchestration” layer. It’s mostly on open weights, by design.
I haven’t yet started SFTing since my examples corpus isn’t quite where I’d like it.
I have solved and formalized a one non-Erdos conjecture. I have formalized several pieces of another subfield that does not exist in Mathlib yet.
As for what I am currently working on, I have an idea I want to build out about how we might think about sieving algebraic structures to generating new, insightful conjectures.
Using LLMs and distributed compute in this context is just a consequence of needing tools to help visualize or materialize things that I am otherwise bad at so I could keep doing the interesting things myself.
> contributions
One thing I think you and the other "AI math builders" have done, is to show how good the top models are at logics and reasoning.
I didn't realize how good they are, until they solved an Erdös problem. And now lots of Erdös problems!
(Plus verified that the AIs actually did solve the problems, that's not easy :- ))
What was really interesting is that during the process it was able to find lemmas or theorems that might be related or relevant to be published.
While I was doing that I was also trying to use Aristotle to do the Lean formalization and I have a WIP system to do that at https://github.com/aconsapart/thesisus/
"He is currently CTO at Xinobi AI, a Japan-based startup developing personal AI agents."
How many of these are you paying for out of pocket??
If you built yourself out of used parts you could do it for under a grand back then too.
(Or millions of disconnected stakeholders with different incentives collectively built modern civilization, but who wants to put that on a bumper sticker)
Are there practical applications of any these problems being solved? No judgement implied, I'm well aware that "no" only means "not yet".
To answer your question directly: most Erdős problems don't have practical applications on their own; the value is the techniques and the machine-checked proofs they leave behind. But there's more real world value in solving some of the FrontierMath Open Problems or Millennium Problems. There's a Venn diagram of "hard problems" and "real world impact" for sure.
The big impact will be with scaling, for example complete autoformalization of existing math, and automatic exploration for new conjectures, with emphasis on how interesting they are. Automatic conjecture generation goes way way back, to the days of Lenat's AM system. Modern AI should do a far better job.
1) For #129 a couple people pointed out that the report was very confusing. And I agree. So I'm currently attempting to improve it.
2) For #130, a person pointed it out as being a partial solution. This seems correct, so I'm currently working on making it fully end to end.
These are put out as "proposed solutions" for the mathematics community to scrutinize, and the scrutiny worked exactly like it should. Happy to take any feedback and make them better.
GPT-5.6 is a closed source model and this seems to be a personal project and not something done by OpenAI.
Unfortunately P vs NP, on the other hand, is going to have to wait for GPT 7
If it were really just about funding people who like math to have fun then it's easy to do forever: just don't have them look at the results and keep paying.
Otherwise they’ll be the ones like Erdős who pose the questions in the first place.
Either way it will always be humans who decide what matters. AI is speaking our languages, not the other way around. We’re in charge. It’s impossible for us not to be, unless we can train an AI from dolphin data or other natural phenomenon.
The AIs intelligence is tuned to us and in 300 years we’ll need new training runs for the update from human zeitgeist language and the 2200 century famous mathematicians.
AI companies are accruing power by virtue of its knowledge and ability to do work. If endowed with agency, which seems likely at this rate, it is the AI itself that will be powerful. And we'll be in charge because AI is trained on human language? I can't fathom the logic behind this.
Put differently: There is no natural law of the universe of why we pay people to do work. It just works well for us currently. If it stops making sense we do something else that works well for that new currently.
There is real risk in doing it too early as well. Imagine if we thought the same ~50 years ago when microcomputers started to appear. If the tech companies had to pay the salaries of every displaced job we would have ended up with severely hampered progress on computers and then a job shortage as we realized there were just different jobs now, not necessarily no work quite yet.
An automatic proof solver doesn't make mathematicians obsolete any more than the excel sheet made accountants obsolete.
Mathematicians will soon be left only to conjecture, with proofs being automated. The issue I see is that AI will devise proofs that are beyond our comprehension, since humans are already taxing each other (cf. Wiles, Mochizuki, Perelman, etc.) Once humans lose grasp of the proof, how will they propose new conjectures?
The connection between “needing to work” and “right to continue existing” is THE problem with society right now.
You think AI boomers are paranoid because they think robots are going to replace the jobs? People are paranoid because they don’t know how they’re going to exist in a post-work world.
There will be fewer good mathematicians, but more great mathematicians!