I have never seen an AI or a human produce a false proof without explicitly using weird meta programming tricks that are very suspicious. No "good faith" Lean proofs have every been shown faulty, to the best of my knowledge. While the risk is non-zero, many of the AI companies are also trying to find bugs in the Lean kernel, so it is becoming very well stress-tested.
would you assess Metamath systems more robust in adversarial settings, because the verifier is so short?
Metamath is short, which does make it easier to verify. In addition, because it's simple, there are many implementions. The set.mm Metamath database, the most popular, is checked by 5 independently implemented proof verifiers.