Say that AI gives you a Lean proof and says it proves Theorem X. It could just as easily give you the same proof but claim that it proves (not X). How would you know the difference?
Nothing can really be considered proven unless a human expert can read the Lean proof and determine that (X as defined in the Lean proof) corresponds to X. The proof (at least the statement of the theorem) must be intelligible to humans to have value.
It's possible people will just start taking AI at its word. Maybe AI says "Here is a Lean proof of X" and we all just shrug and go "Okay, X is proven." But that's not how it works right now for human mathematicians. Why would we apply that standard for AI?
The statement of the theorem has to be correctly translated from English into Lean code.
It’s like translating user requirements into code. The code could run without bugs but not do what the users want.
The only way to know the AI did it correctly is to check. You can’t just take it at face value.
The "magic" of lean is that (in principle, assuming lean is sound and the proof is verified) that is all you have to check by hand. That is a big deal.
Edit: Yup. A bug report to Lean was disguised as a "Collatz" proof in a humorous way. Links below.