If they can’t, how do we make them so they can?
Plenty of humans seem to have no problem admitting when they don’t know something.
If they can’t, how do we make them so they can?
Plenty of humans seem to have no problem admitting when they don’t know something.
If that’s true then this behavior difference (identifying and admitting when you don’t know something) seems like a pretty fundamental difference between humans and LLMs at least at this point.
I wonder -- does that kind of interaction come up in spoken conversation more than in written text? Things like blog posts and tweets are often stated very confidently. If LLMs are trained mostly from those sorts of sources, maybe that's why they end up artificially overconfident.
https://www.gally.net/temp/20231015gpt4hedging/index.html
The only confabulations are when reading text. In my tests so far, GPT-4 is not a good OCR engine and does make up a lot.
The samples I posted above don’t contain any major hallucinations about nontext image components, but other tests I’ve done with GPT-4 did. But GPT-4’s image recognition still seems better than Bard’s, which has been pretty bad in the tests I’ve done with it.
GPT-4 logits calibration pre RLHF - [https://imgur.com/a/3gYel9r](https://imgur.com/a/3gYel9r)
Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback - [https://arxiv.org/abs/2305.14975](https://arxiv.org/abs/2305...
Teaching Models to Express Their Uncertainty in Words - [https://arxiv.org/abs/2205.14334](https://arxiv.org/abs/2205...
Language Models (Mostly) Know What They Know - [https://arxiv.org/abs/2207.05221](https://arxiv.org/abs/2207...
TLDR; Hallucinations aren't a randomness or representation issue. The computation already knows. It just doesn't care about telling you this.
But why? Why would the LLM creators want the LLM giving out knowingly bad information? This is what I don't understand.
It makes me not want to use the product. Same as I don't like asking people things who I know like to just spew bad information when they don't know rather than simply saying they don't know.
So how do we get LLMs to admit when they don't know something rather than confabulate? Seems like that would be a really great improvement to me.
>So how do we get LLMs to admit when they don't know something rather than confabulate?
The prompts here demonstrably help https://arxiv.org/abs/2305.14975
But otherwise nobody is certain how to do it. The simplest solution is just making them better. Hallucination is something that happens when knowledge/memory fail. By making them more competent, they hallucinate less. So even if they never care about communicating this, you wouldn't notice.