A simple example is to ask it to write a complex program in a language like Java or Kotlin that requires you to declare all your imports up front. There can be a lot of these. Therefore the model must figure out what it wants to use very far down the program in order to emit the first words of the program correctly. Presumably somewhere in the intermediate state process the network is actually working out the entire program ahead of time so it can predict the next word of it, and that's sufficiently stable for it to do so successfully.
Also I'm not quite sure I understand what you mean by not maintaining hidden state between tokens. Doesn't the transformer algorithm start by parallel processing every token of the prompt to compute a large set of vectors and matrices that are then reused over and over whilst generating the results? What are the KV vectors if not hidden state of some sort?