Memory for chatbot is usually in the form of an external memory; aka not compressed into the weights of a neural network; but presented simultaneously as input and selected with some form of attention like transformers do.
In your specific chatbot use case the obvious external memory to present to the bot is the full chat history. Alternatively you can manually extract features from the full chat history and present those instead.
Language models are just always trying to predict the next character.
Language models are kind of the jack of all trades. They are generic and given enough parameters, enough data, enough compute they will learn to solve all tasks simultaneously. But they are not enticed to solve the task you are interested in, they are just trying to learn how to best predict the next character with the finite capacity they have, and will only learn memory insofar it helps them achieve their task.
If you are interested in a specific task you can inject your specific desires either into the structure of the model, or in the structure of the dataset.
There are various line of thoughts when you want to achieve a specific task.
-The GPT approach is to make the model bigger and bigger and solve task as one or zero shot learning. By not touching at the structure of the model anymore but by training it on tasks encoded as structured text.
-You can find some literature about question answering tasks.
You can have a transformer "encode" a big document and answer the question by extracting just the relevant information.
-This mean for example if you want for your chatbot to remember your name or any information you gave him before, one quick way is to give it your whole chat history (it's probably not so big). It's akin making the context window big. You can use tricks like LSH tranformers, or some form of hierarchical memory to not be too much memory constrained by the attention.
-You can also encode "manually" : Train a separate neural network to answer the question you are interested about from the chat history. Like "what is his name ?", "how old is he ?",... And save it as a context information that will be presented as input knowledge for your specific chatbot training task.
-You can also update the weights continuously like it is done in reinforcement learning. Fixed-sized neural networks need to be presented the information multiple times before being able to ingest it. So you'll need to use some kind of replay memory. It's also not a great idea to update all billions of parameters every time you want to retain an information so you will probably need to use sparse operations such that only a few parameters are updated at a time. If you look hard enough you'll probably notice that most traditional database operations can be encoded as sparse neural network operations. If you look even harder you'll probably notice that a lot of information retrieval algorithms are just gradient descent on some form of sparsely encoded neural network operations.
-You can also use GAN for text so that the loss function optimized is closer to the task at hand.