I’m struggling to understand what’s novel here. LLM-based summarization of chat history memory is a well-established technique implemented by many LLM frameworks. Summarizing on every message is, as proposed in the paper, a major performance bottleneck and adds significant latency to the chat loop.
Many implementations utilize a fixed sized buffer, progressively summarizing batches of older memories when they fall out of the buffer. Ideally, this is also done out of band to the chat loop.
I’m an author of Zep[0], an open source long-term memory store, and this is how we implemented summarization.