I was designing a system that pulls structured data out of an unstructured transcript, and Opus initially suggested a confidence score for each value. When I asked it “do you think the model will generate a meaningful number there?” the answer was basically “ha, you caught me.”
What did work was asking the model to emit citations in the form of line numbers from the transcript. It dramatically dropped hallucination.