SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationales
arxiv.org
arxiv.org
https://arxiv.org/abs/2207.05221
"Calibration" in a knowledge context means having estimated_p(correct) ~ p(correct), and it turns out that LLMs are reasonably good at this. Also a core reason why LLM-as-a-judge works so well: quality evaluation is vastly easier than generation.
This is one of the key prompting in a lot of Enterprise cases. You can currently prompt LLMs to add a confidence score along with their responses.
Especially when you are using LLMs for downstream NLP tasks.
The confidence score can be a great indicator also for applying a two-tier model approach!