The LLM was happy to give you 'happy' but incorrect answers because it thought it'd make you the most satisfied. Kinda like a psychopath car salesman who just wants to make a sale.
The LLM was happy to give you 'happy' but incorrect answers because it thought it'd make you the most satisfied. Kinda like a psychopath car salesman who just wants to make a sale.
Do LLMs "know" that they don't know?
Real understanding includes acknowledging what you don't know.
Can the predictive text generation of LLMs recognize that its training set did not include data?
GPT-4 logits calibration pre RLHF - https://imgur.com/a/3gYel9r
Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback - https://arxiv.org/abs/2305.14975
Teaching Models to Express Their Uncertainty in Words - https://arxiv.org/abs/2205.14334
Language Models (Mostly) Know What They Know - https://arxiv.org/abs/2207.05221