Having a LLM recall something with exact detail some 100k tokens ago sounds a bit like the ADHD test Cartman got in South Park. We don't recall exactly but rather a summarized version.
On the other hand, computers recall exactly when asked directly (RAM access) so in that sense it seems natural to want that from a LLM.
One thing we can do which current LLMs can't, at least directly as far as I'm aware, is to go back and re-read a section. Like on-demand RAG, or something.
In the meantime, good to know it's not terribly useful to have the full 128k context, as it usually is too much for my GPU anyway.
Encoders can do that. And we can use them with diffusion to generate text [0].
This works because you don't impose a masked self attention for autoregressive decoding in the encoder, so subsequent layers can re-focus their key/query vector space to steer "backwards" information flow.
Happy reading. Feel free to get back!