GPT is auto regressive. That means each output token becomes part of the new input sequence. Which is to say, the beginning of the model’s answer becomes part of your prompt.
If the model makes some mistake in the beginning, it now needs to explain / make sense of that mistake.
Kind of like a split-brain patient whom you ask why they got up, and they then say, to get a Coke. [1] In psychology, that is called confabulation. In machine learning, they use “hallucination“, probably so they can use the term across several disciplines, like language, audio, vision, etc.
[1] https://www.brainscape.com/flashcards/chapter-4-hemispheric-...