> Which makes sense to me. If an LLM is giving you a mixture of "hallucinations" and correct answers, the correct answers will similar and the hallucinations will hopefully be chaotic
I expect that to give you something close to the confidence of the underlying model to some specific claim, which is good, but I still expect legends (urban and cultural) to be high-ranked.
They'd be very human mistakes, but still mistakes.
I think the only way past that is to build a world model, look for contradictions, and then look for new evidence to resolve those contradictions.