If I ask you to envision a green triangle and a red square next to each other, and then swap the shapes but keep the colors in the same locations, and answer what color is the triangle now, you say the triangle is red, but you do so because you envisioned the triangle swapping places and did the mental steps etc.
An LLM if even answering correctly, is statistically answering based on billions of lines of text + rlhf and all of this, I highly doubt there is a mental model of the world, but rather a large set of constraints in the probabilities which leads to the resulting answer. The reasoning ability is a secondary effect of the probabilities which is why it's hard to make it so every probability is correct for every answer I think.
And regarding OP about hallucinating vs confabulating. To me hallucinating is a fine word for it, because it is filling in a gap or there aren't enough constraints in the model/data/tuning to account for that specific answer that it gave that was incorrect. Hence it "hallucinates" something in the gap. The real power of LLM's is that it seems to accumulate these 'constraints' (generalization), so that with the right model, it should be able to answer more and more prompts correctly, which is kind of amazing.
Confabulation works too but is a little more high level IMO.