This is just ticking a checkbox. And even worse and more time wasting even: you have ZERO proof that there's no latent Lean bug. Especially in a proof this large.
Maybe. But they're in 32 millions lines of Lean. How do you find the needle in the haystack ?
>Or, you could prompt agents later to analyze which lemmas or parts of the proof are surprising or applicable to other problems?
If OpenAI was truly serious about improving maths (and not jerking themselves off), they'd have also used Prove2Me (and contributed their results back), which would have done that. Each part of the proof combines into a larger graph, that everyone can reuse. Note that Anthropic isn't better there: yes, they used Prove2Me, but as far as I know they haven't contributed back to it, and just shat out 10 million lines and a good luck everyone.
Is there an established term for the idea of "DoS"? I've taken to calling it slop fatigue.