Its not rationalising, its statistically choosing the next token. Also, generally the LLM doesn't get defensive (although its not always the case.)
Its not like a dementia patient not finding their keys and gradually convincing themselves that someone stole them instead. Because the alternative, that they are loosing cognitive faculties, is too horrific to acknowledge.
This is why its so annoyingly problematic. too much humanising, and not enough witnessing someone with dementia "confabulating"
It is noise, and nothing else.
We know how the LLMs were trained, so re-stating it doesn't help in anything. The point is that after the LLMs are trained they behave in certain ways and it can be helpful to say something about how it behaves.
For example, we can talk about how a linear regression can or cannot capture the causal effect of X1 over Y. "It's not capturing the causal effect. It's just minimizing the squared error" is an unhelpful statement.