Consider this: while the inner state of an LLM (all its activations, residuals stream that is cached in the KV cache) is fully deterministic given its input sequence, the information contained in it IS NOT identical to the information in the input sequence. The reason is obvious: the LLM itself contains an enormous amount of information in its parameters and it transfers it to its residuals stream at each forward pass.
In other words: the final state given the two input sequences (where NT stands for "null token"):
<problem-prompt> [NT]
and
<problem-prompt> [NT] [NT] [NT] [NT] [NT] [NT] [NT] [NT]
is not the same, and at each forward pass the LLM keeps working on the solution even if the input tokens provide absolutely no further information.
If this is correct, then there is no need for the model to have already verbalized the key elements that drive the "aha" moment, so no need for the "aha" to appear after a full explanation.