Things that the software running the model would otherwise recompute, if not for the cache. What special meaning are you assigning to it?
From what I understand, at position Aha in each layer it's constructing a query based on the current activation and looking at the key of each other token position for that layer, in order to decide how much attention to pay to the value.
In this way it attends to the previous values, such as perhaps the incorrect assumption and plausible explanation.