All these 2k/4k/8k context sizes that we've had recently are able to map pretty well to what a human could reasonably remember. What I mean is, you could ask a human to read some text with 8k tokens, and for the most part they could answer questions about the text coherently.
But what about 32k contexts, or beyond? At some point, as token size increases, the ability of a human to give a highly precise and detailed answer decreases. They must start to generalize. A human could not read Infinite Jest in one pass and then answer details about every single sentence. But could a transformer, or a RNN? As the context grows, is it harder to keep a high granularity of detail? Or am I wrong in trying to think of these models the way I think about the human mind, and they are actually able to handle this problem just fine?
I'm aware that we can cheat a bit, by adding a lookup step into an embedding database, to provide "infinite" context with "infinite" precision. But to me, that is analogous to a human looking up information in a library in order to answer a question. I'm interested in the inherent, emergent memory that these models have available to them in just one forward pass.