https://github.com/openai/chatgpt-retrieval-plugin#memory-fe...
> how to work with a memory module that remembers things about specific entities. It extracts information on entities (using LLMs) and builds up its knowledge about that entity over time (also using LLMs).
[1] https://python.langchain.com/en/latest/modules/memory/types/...
An alternative could be a vector store, injecting small snippets of relative text as a step.
0 - https://python.langchain.com/en/latest/modules/memory/key_co...
"We show that transformer-based large language models are computationally universal when augmented with an external memory. Any deterministic language model that conditions on strings of bounded length is equivalent to a finite automaton, hence computationally limited. However, augmenting such models with a read-write memory creates the possibility of processing arbitrarily large inputs and, potentially, simulating any algorithm."
From "Memory Augmented Large Language Models are Computationally Universal"
https://deepai.org/publication/memory-augmented-large-langua...
I suspect that if you filled the context window with "1 1 1 1 1 1 1 1 1 1", and then asked "How many 1's did I just show you?", it probably wouldn't know, simply because whatever tricks they use to have such an apparently large context window don't allow it to 'see' all of it at any given moment.
So turning a 4k window to a 32k window means a 512x increase in compute they'd need (just to maintain similar output quality).
I suspect they must have found a better solution to be able to scale the window so big. They haven't announced what it is.