There's far more than zero evidence, see https://arxiv.org/html/2402.08806v1 ; or the industry-standard practice of using multiple model families as LLM judges; or even 1P implementations https://code.claude.com/docs/en/advisor (which misses most of the benefit; since you want different model families, with different pretrains and posttrains).