Neither does GPT. The whole conversation is fed back to it everytime. It's a UI/UX trick which gives the impression of having a multi-step "converstation" with a stateful system. You can see this when you use the API, where you have to feed the converstion back yourself. This can be replicated in ChatLLaMA.
* inject the running chat log into the prompt * inject the summary of the chat into the prompt
You can also fine tune the model to incorporate larger amounts of data, but that may be more expensive (and slower)
This kind of sounds like human short term and long term memory. Maybe “fine tuning” is analogous what happens to our memory when we sleep.
Alternatives are maybe architectures using langchain or toolformer to retrieve "memories" from a database by smart fuzzy search. But that's worse, because reasoning would only be done on that context, instead of all memories it ever had.