ChatGPT Isn't 'Hallucinating'–It's Bullshitting – Scientific American
scientificamerican.com
scientificamerican.com
The term "hallucination" may be imprecise, but it is a lot closer to the truth than "bullshit".
OTOH, "hallucination" is defined by Cambridge [0] as:
> the experience of seeing, hearing, feeling, or smelling something that does not exist
That can clearly not be the case for an LLM as it lacks the senses to see, hear, feel or smell. Arguably, some systems can see and hear but I don't think they would be able to sense anything that does not exist.
The dictionary follows on by including LLM hallucinations in its definition, but this seems irrelevant since that's exactly what's in dispute in this post.
[0] https://dictionary.cambridge.org/dictionary/english/hallucin...
https://www2.csudh.edu/ccauthen/576f12/frankfurt__harry_-_on...
It still seems like the term in its 1986 definition is a good fit for what we currently use Hallucination to describe.
The bullshit is only in the actions of marketers trying to engage technically illiterate wallets in funding a bubble out of false promises (such as "dreaming", "imagination", "hallucinations", "bullshitting", "superintelligence", "superalignment", "smarter than human", and so forth).
This is just wrong. Accurate modelling of language at the scale of modern LLMs requires these models to develop rich world models during pretraining, which also requires distinguishing facts from fiction. This is why bullshitting happens less with better, bigger models: the simple answer is that they just know more about the world, and can also fill in the gaps more efficiently.
We have empirical evidence here: it's even possible to peek into a model to check whether the model 'thinks' what it's saying is true or not. From “Discovering Latent Knowledge in Language Models Without Supervision” (2022) [1]:
> Specifically, we introduce a method for accurately answering yes-no questions given only unlabeled model activations. It works by finding a direction in activation space that satisfies logical consistency properties, such as that a statement and its negation have opposite truth values. (...) We also find that it cuts prompt sensitivity in half and continues to maintain high accuracy even when models are prompted to generate incorrect answers. Our results provide an initial step toward discovering what language models know, distinct from what they say, even when we don't have access to explicit ground truth labels.
So when a model is asked to generate an answer it knows is incorrect, it's internal state still tracks the truth value of the statements. This doesn't mean the model can't be wrong about what it thinks is true (or that it won't try to fill in the gaps incorrectly, essentially bullshitting), but it does mean that the world models are sensitive to truth.
More broadly, we do know these models have rich internal representations, and have started learning how to read them. See for example “Language Models Represent Space and Time” (Wes & Tegmark, 2023) [2]:
> We discover that LLMs learn linear representations of space and time across multiple scales. These representations are robust to prompting variations and unified across different entity types (e.g. cities and landmarks). In addition, we identify individual "space neurons" and "time neurons" that reliably encode spatial and temporal coordinates. While further investigation is needed, our results suggest modern LLMs learn rich spatiotemporal representations of the real world and possess basic ingredients of a world model.
For anyone curious, I can recommend the Othello-GPT paper as a good introduction to this problem (“Do Large Language Models learn world models or just surface statistics?”) [3].
[1]: https://arxiv.org/abs/2310.02207
This makes sense, since you would expect LLMs to perform better when they can differentiate falsehoods from truths, as it's necessary for some contextual prediction tasks (say, the task of predicting Snopes.com, or predicting what would a domain expert say about topic X).
No. They are functions of their training data. There is absolutely no part of a LLM that functions as a truth oracle.
If training data contains multiple conflicting perspectives on a topic, the LLM has a limited ability to recognize that a disagreement is present and what types of entities are more likely to adopt which side. That is what those studies are reflecting.
That is, emphatically, a very differing thing than "truth."
Again, we have empirical evidence to suggest otherwise. It's not that there's an oracle, but that the LLM does internally differentiate between facts it has stored as simple truth vs. misconceptions vs. fiction.
This becomes obvious by interacting with popular LLMs; they can produce decent essays explaining different perspectives on various issues, and it makes total sense that they can because if you need to predict tokens on the internet, you better be able to take on different perspectives.
Hell, we can even intervene these internal mechanisms to elicit true answers from a model, in contexts where you would otherwise expect the LLM to output a misconception. To quote a recent paper, "Our findings suggest that LLMs may have an internal representation of the likelihood of something being true, even as they produce falsehoods on the surface" [1], and this matches the rest of the interpretability literature on the topic.