Again, you don't need to check the proof, Lean does that. You need to check the statement, which is a much easier task. So yes, you are wrong. Skepticism is good, but it is usually just ignorance.
Yes, you are misunderstanding what the paper claims. Navier Stokes for example is not such an example, only the proof is different. For their other examples, none of them they claim that they concern actually the open ai solved theorems. You are welcome.
Yes, Grant Sanderson might be a great example. I am not doubting the competence of the "good lecturers" I was referring to. But that is different from being able to introduce the big new concepts that give the big new results.
No. What the paper says is that in principle, translating NL statements to Lean statements is hard. Nobody doubts that, translating informal to formal text cannot be formally proven correct, so...
Does the paper give a single example of one of the OpenAI solved theorems with a Lean certificate where the Lean statement does not correspond to the actual statement from the mathematical literature? I don't think so, but in case I am wrong, feel free to provide that example.
Jesus Christ, so many people here who have no clue what they are talking about.
A proof of a theorem is different from the statement of the theorem. OpenAI has a Lean proof of the statement. That is all they need. There may be many different proofs of this statement, including NL proofs. It does not matter that these NL proofs may or may not be different from the Lean proof, at least for the correctness of the Lean proof. But of course the NL proof may be wrong. But who cares?
What I am seeing is not people not being diligent in their research, but people not being diligent in their reviewing of research. But maybe there is a positive feedback loop here...
We all run into hallucinations like in your example, but that doesn't mean that the AI cannot make connections between different areas of knowledge it has, and combine them usefully to something new that wasn't there before. Like Navier Stokes, for example. I've had it make plenty of such connections in my own research, giving me new results I didn't have before.
We don't lose progress by AI solving Navier Stokes. Get a grip on yourselves. I don't think there is a "progress" argument here without tying mathematics to "usefulness", and AI makes mathematics dramatically more useful.
Just like my AI. They also think before and during a code refactoring. There is a gap between my prompt and them taking actions, you could call that "stopping and thinking".
Maybe tiring, but that is the reality. Note that mathematicians are only upset now that "have you tried this on the latest model" works for so many of their problems now, but didn't for the model before that.
When encountering a difficult problem, add an indirection.
That indirection, in the case of LLMs, is formal proof. It can actually turn an LLM into a sort of compiler. Where, if the compiler run completes successfully, you don't need another run, and you are sure it is correct.
As someone with a math background, and a logic background, I think Russell's paradox is most interesting, beyond just set theory. It shows us that one should be careful about what exactly a definition is.
If you see any post on arXiv you should assume it is junk.
I find it ridiculous that people put any value on something being posted on arXiv. That doesn't mean the post is bad. It just means you need to find other means of judging it, for example by actually reading it.
> We are currently facing the very specific challenge of advising OpenAI on how to coordinate the release of a large number of significant results in mathematics that they report have been produced by their internal model.
The only thing I want to see is the problem statements, solutions, and associated Lean proofs. Anything else is gatekeeping. What a low point for academia.
Intelligence does not seem to be magic either, as LLMs are proving now. It is indeed a waste of time to argue that LLMs are not intelligent in their own way, they obviously are. If Navier Stokes doesn't convince you, nothing will.
I just used a £89 Codex subscription to do very intelligent things with it, stuff that I would have had to sit down and ponder and work on for quite a while, and I have a PhD in that. I didn't need to do anything special except explaining the problem(s) to the AI, and my theory of it so far. It took it from there. If that is not intelligence, nothing is.
I think this axiom is of course true. But the mistake the article makes, in my opinion, is to try to apply this axiom separately to each domain. If we have this as the over-arching axiom, it is not clear at all that humans should be steering the development of mathematics. Maybe it would be better for humanity if the department of world math is run by AI.
I am shocked how people can deny that solving Navier Stokes requires some sort of intelligence. Even Doctorow talks of "brute-forcing" a solution. Brute-forcing leads to combinatorial explosion, so there must be something more going on here. Otherwise you could just put this problem into an automated theorem prover (we've had those forever, they are actually just brute-forcing it).
I wouldn't use SPARK either, and rather develop my own approach. The problem isn't that the proof has 400 lines of code, every modern system has large proofs (Isabelle/HOL, Lean, etc.) My latest formal proof has over 50K lines of proof. That's why AI is such a useful tool.