The problem is how do you know whether the answer is just the most persuasive or actually the most accurate one? It's hard to figure this out without domain knowledge.
I do something similar with reviewing code: I have one agent write the code and another reviews it, then they go back and forth for a bit improving the code. Seems to yield better results than one agent alone.
Seems like a similar principle.
https://www.nature.com/articles/s41746-026-02619-0
https://www.nature.com/articles/s44360-025-00007-8?fromPaywa...
Different prompt approaches and training doctors to use LLMs can improve accuracy of LLM-assisted diagnosis. It’s pretty reasonable to hypothesize that LLM “peer review” could improve that as well.
I never said it can't work. I just said that finding the correct medical digagnosis is different than finding a solution to a software problem.