It feels lazy that this isn't built into the harness in some adversarial citation review checkbox.
Also the questions LLMs ask are so inhuman. It's like it's some contrarian trying to win an argument over getting an answer.
BUT
On models with web search, they can scan the sources in context, and those are pretty good at correct citations. However it's still a "trust, but verify" situation.
It is commonly this way right now, but it doesn't have to stay this way forever.
It is also common, in my usage at least, to iteratively brow-beat the bot into paring its statements down to those that which are supportable by its sources. Doing so just takes repetition, and that repetition takes time and burns more tokens.
With the present state of things, the prompts to get moving on this and to guide the ultimate response into something that is verifiably supportable by outside sources can usually be simple and largely generic.
They're easy enough prompts that a subagent can produce them.
(I've done it myself with Codex subagents and it worked very well, aside from the unsustainable burn rate that did not fit my budget.)
Think about your coding harness: the model reads a bunch of files into the context, and it generally doesn’t forget/hallucinate which lines came from which file.
I mean, its not. Lots of LLM's focused on finance do this already.
Bloomberg's own ASKB produces results and provides links back to the source documents or urls that it references so people who care about correctness can verify the results.