I describe it as more of a “mashup”, like an interpolation of statistically related output that was in the training data.
The thinking was in the minds of the people that created the tons of content used for training, and from the view of information theory there is enough redundancy in the content to recover much of the intent statistically. But some intent is harder to extract just from example content.
So when generating statically similar output, the statistical model can miss the hidden rules that were a part of the thinking that went into the content that was used for training.