Oh interesting, didn't know. How does this work past the first transformer in the stack?
This means that every attention layer can use previously calculated outputs for the same prompt prefix. So it only needs to calculate from scratch starting from the first unique token in the prompt sequence.