Let's say we want to let our chat remember the character slammed the door last time they were in Village X with the mayor in their presence and have the mayor comment next time they see the player.
Every X tokens we can fire a prompt with a chunk of conversation and a list of semantically similar entities that already exist, letting the LLM return an edited list along the lines of:
entity: mayor
location: village X
priority: HIGH
keywords: town hall, interact, talk
"memory, likelyEffect"[]: door slammed in face, anger at player
Now we have:- multiple fields for similarity search
- an easy way to manage evictions (sweep up lowest priority)
- most importantly: we're providing guidance for the LLM to help it ignore irrelevant context
When the user goes back to village X we can fetch entities in village X and whittle that list down based on priority and similarly to the user prompt.
None of this has any determinism: instead you're optimizing for the illusion of continuity and trading off predictability.
You're aiming for players being shocked that next time they talk to the mayor he's already upset with them, and if they ask why he can reply intelligently.
And to my original point while this works for a game-like experience, you wouldn't want to play around with this kind of fuzzy setup for your companies internal CRM bot or something. You're optimizing for the exact value proposition of your use-case rather than just trying to throw a raw RAG setup at it