There's far more than zero evidence, see https://arxiv.org/html/2402.08806v1 ; or the industry-standard practice of using multiple model families as LLM judges; or even 1P implementations https://code.claude.com/docs/en/advisor (which misses most of the benefit; since you want different model families, with different pretrains and posttrains).
Eh that’s about medical diagnoses. Not directly transferable. For regular usage, multiple models just make you feel productive but I bet they aren’t any more productive than just one.