Why truncate the conversation? Why not have the LLM generate a summary, and use that instead of a truncated conversation?
That seems closer to what we do, we recall the important bits without every single word said.
That seems closer to what we do, we recall the important bits without every single word said.
But its also several passes through the LLM instead of one, and is prone to distortion.