You are right on fhe first point, I believe the second needs to be corrected. LLMs have no state. Each word- each token- is made anew, the network being oblivious to the process that led to the previous token. It cannot even distinguish between text it generated and text you added unless you tell it that. So any relationship between output and explanation is 'coincidental'. Actually, any relationship between
two tokens next to each other is 'coincidental' in the exact same way, so it cannot explain its reasoning. If you prompt it to comply or explain why not, the network may generate an explanation of why it can't comply, but it is doing that simply as a good continuation of the input text.
I'm not saying you can't call that an explanation if you want to use that word, I'm trying to make a distinction between the generated output and an explanation of that output. All of this is the former. The latter is not possible unless you analyse the model as it is running, and we don't really have a good understanding of that process yet.
There was a cool paper where researchers analysed a finetuned model playing othello, and they found some neurons that clearly represented the game state- if they changed those neuron values, the network responded correctly, even if that game position was impossible to get to normally [1]
I'd argue work like that is getting us closer to getting explanations for llm behaviour.
[1] https://www.joyk.com/dig/detail/1674380676520322