Either way this type of idea will probably be a fundamental feature for all chat bots in the future IMO.
Either way this type of idea will probably be a fundamental feature for all chat bots in the future IMO.
From what I remember the title in ChatGPT gets set once after a few messages, in which case I'd assume it's generated with a special "title generation" prompt (that gets the first few messages as input).
In either case since I don't work at OpenAI I can't tell you for sure ;)
There are trivial techniques to implement "lossy" memory, such as just average pooling tokens (the same approach used by sentence transformers). Not sure why it's so rare to see this used for condensing a huge amount of context into a prompt. It is effectively "medium" term memory.
Fed chatGPT special numbers, then 3k tokens, then 2k tokens. after that, it was unable to understand any question about the special numbers provided.
https://chat.openai.com/share/8a0675b6-2876-4606-ac79-646391...
The problem could be some kind of instability of attention as it scales above 10k tokens. A recent paper suggests attention mechanism needs a default value (a "sink"), and its absence produces instability.
https://arxiv.org/abs/2309.17453
Another paper says the middle part is lossy while the beginning and end are better attended.
I have also been dabbling with neural nets (pre-transformer), especially LSTM which have a "residual" connection, the one I was mentioning. That makes gradients better behaved. Schmidhuber tech.
https://www.kaggle.com/code/sid321axn/regularization-techniq...