This seems to be a fundamental side effect of next token thinking, training data, etc.
This seems to be a fundamental side effect of next token thinking, training data, etc.
There are solutions, but no quick band-aid.
You could RLHF/whatever models on common factual questions to try to get them to answer those specific questions better, but I doubt there'd be much benefit outside of those specific questions.
There's a couple of fundamental problems related to factuality.
1) They don't know the sources, and source reliability, of their training data.
2) At inference time all they care about is word probabilities, with factuality only coming into it tangentially as a matter of context (e.g. factual continuations are more probable in a factual context, not in a fantasy context). They don't have any innate desire to generate factual responses, and don't introspect if what they are generating is factual (but that would be easy to fix).
Or maybe this is inherent to continuation?
The behavior reminds me of the human subconscious, which doesn't say no, just raises up what it can.