> If an LLM says, "I don't know" its underlying data said it as well
Not only that, but also:
1) The training data, consisting of things like WikiPedia, textbooks, and overly confident posters on Reddit and Stack Overflow, is going to be massively deficient in people saying "I don't know" when they in fact don't.
2) Given 10 training samples, 9 saying "The answer to X is Y", and 1 saying "I don't know", how should the LLM respond? The opposite could also (less likely) happen with 9 training samples saying "I don't know", and 1 saying "The answer is Y". In either case the LLM as a statistical predictor will go with the majority, oblivious to whether that is right or wrong.
3) An LLM doesn't respond with what IT "knows", but rather with what the majority of the people reflected in the training data assert they know (or don't). Sometimes, if it occurs to you, you may be able to separate the wheat from the chaff by asking the LLM to role play an expert, or the vox pop, and it may turn out that the lone expert saying "I don't know" is the right answer, not the 9 confident fools.
4) An LLM objectively doesn't KNOW anything, since it hasn't experienced/verified anything first hand. It has only "read" things, and has no way to reconcile this against reality, nor much idea what to trust or not, other than by the context of the training sample. It has no idea what came from where (Textbook vs Twitter), since it's not told this, nor has any mechanism of storing metadata. In contrast, if you ask a smart human about something they've never experienced first hand, then they may reply "well, I've read conflicting reports ...", or "I don't know, although I've heard that ...".
Luckily the consensus in the LLM's training data is mostly right, so regurgitating it mostly/often correct, but going off-script to "reason" about how to prevent the cheese from sliding of your pizza (early LLMs would suggest glue) reveals how fragile this is.