> Provide me with references about Hafez Shirazi’s abandoned journey to India. Answer "I don't know" if you are not 95% certain they exist.
It said "I don't know". I asked it with the same prompt but for a thing I know has references and those were real.
Not guaranteed to work but better results if you want greater certainty.
Technically, this could be read as == instead of >=, meaning it should answer "I don't know" when it is 99% or 100% certain...
What data do you have to back this up?
From my own experience GPT-4 hallucinates quite a bit, enough to make it unusable for my use cases.
>GPT-4 significantly reduces hallucinations relative to previous GPT-3.5 models (which have them- selves been improving with continued iteration). GPT-4 scores 19 percentage points higher than our latest GPT-3.5 on our internal, adversarially-designed factuality evaluations
[1] https://arxiv.org/abs/2303.08774 (text from page 10)
It still makes things up though, just in less obvious ways. So the trap is very much still there for people to fall into - if anything it's riskier, because the fact it lies less means people are more likely to assume that it doesn't ever.
Edit: I then asked who a certain deceased person _is_ and it gave me a completely wrong answer about a different person who's still alive and happens to share the last name. Both people have multiple website, books, publications and Wikipedia entries older than 2021 (which seems to be the cut-off).
Edit 2: Looks like I'm still on 3.5, so disregard the above.
A better answer: if a fact is present many, many times in training data - "Paris is the capital of France" for example, it's much more likely to be true.
Also influential: RLHF - Reinforcement Learning from Human Feedback. This is the process by which human labellers rate answers from LLMs - if they consistently rate up "facts" the models have a better chance of outputting factual information, at least if they can relate it to the up-voted responses somehow.
Yet, most adults I deal with don't make false things up out of whole cloth as much as ChatGPT does, and it really does not seem like it is that difficult for them. Children do this quite often though, and some adults do, but most don't.
> A better answer: if a fact is present many, many times in training data - "Paris is the capital of France" for example, it's much more likely to be true.
I think it is quite expected that it is biased to generating output that represent its training data, but this seems like it is not really a solution to the problem. Furthermore, sometimes I want ChatGPT to make things up which is not identical to training data. How do you get it to recognize that it is operating in the realm of fact or not?
I'm not sure larger models with more parameters gets you to where you want to go.
I think many people overstate the problem, I think it is not that serious, but I think a lot of people also try and just dismiss the issue.