This is not how science and engineering work and an arxiv should not be taken at face value.
Simply asking the LLM in two separate contexts the same question but from opposing perspectives, then in a third context asking it to analyze both responses and choose the most neutral and objective take, you wipe out any "(dis)agreeableness" bias and dig closer to a deeper, more nuanced synthesis of a given topic. This paper is just taking this idea to the next level.
This isn't really possible with RLHF alone unless you train the LLM to often give two opposing perspectives, which would get tiring.