How do you know that it's formalizing what you think it's formalizing? If your Lean 4 has a bug, won't you be proving something other than what you thought?
These axioms don’t have to be the core axioms of math. If some other result has been formally proven, I presume you can simply use that result as an axiom.
As long as you do those things, what happens in between is immaterial from a correctness point of view because each of those statements is proved by the statements before them.
You also have to check for things like sorry or defining axioms.
The computer could generate a huge document, how would you check that it's right?
I don't get what the controversy is about. Are people expecting AI to be perfect? Do they think they won't have to do the work to verify it themselves?