Let's say that the forward pass that selected "Aha" produces activations that indicate a wrong assumption, and a plausible explanation.
It puts learned projections of the activation into the KV Cache and outputs Aha.
Both the cached projections and the current Aha token can now influence further activations in an additional Forward pass that the Aha bought the model.
At least that's how I thought it works.