That rests on a false-assumption that the errors are statistically independent events, and have nothing to do with the shared nature of the judges.
And even if they are related - if Opus 4.8 always has a 1:100 chance of a specific hallucination - then running the same model twice does indeed dramatically reduce the odds of an error in the final output.
Yes, LLMs can be SOTA for NLP, but you’re going to have to use them to write software or workflows that are more deterministic.
When they don’t know something, they figure it out empirically. For things they already know, they are consistently correct.