I don't know but I would expect it to be realtively easy for an LLM to detect "hallucinations".
I don't know but I would expect it to be realtively easy for an LLM to detect "hallucinations".
I think this may be part of the problem. The actual humans creating the report don't have the expertise to know which one to trust. At least that was what consulting was like in my experience at a similar firm.
Yes, this technique and its variations[1][2] "work" but it's still not 100% perfect. And it's not as widely used it might be because, among other reason:
a. it takes longer to implement
b. it costs more (more tokens spread across multiple llm calls)
c. higher latency (getting an answer takes longer due to multiple llm calls involved)
d. the final answer is probabilistically more likely to be correct, but is still not guaranteed to be error free, so you can never fully escape the need for Human in the Loop.
Then, it doesn't matter if you add 1000 frontier models -- they still can't generate a good report.
But yes I suppose you can get rid of hallucinated citations though