On the other hand, if a proof exploits a bug in Lean then it's ~trivial to prove a contradiction, so you can just check all the proof steps to see if it also allows doing that.
Similarly if the formal Lean problem statements (human generated) are correct translations into Lean (and the original NL statements are sound, which one would hope after decades), and no Lean bugs are abused by the proof (as defined above), then the proof is valid.
The NL/Lean discrepancies are super annoying and will make human analysis hard and fraught, but as many posters have found out the models themselves will gladly pick apart the NL-Lean translation for errors, and so my guess is that finding the discrepancies will not take too long. OpenAI really should have done a dynamic workflow over every lemma and step to ensure pointwise accuracy in the translation.