Then let it go, it'll continue carrying through and even sometimes updating its 'memory'. (not updating as much as I'd like: it's extremely efficient at just copying text from recent history).
I think it would be interesting to train GPT3 specifically to work this way: First train GPT3. Then, run back over the training data and use GPT3 to generate running summaries: At each paragraph break add some text that says something like "The most important things about the above text is:" and let it complete that prompt).
Then use those running summaries to augment the training data with special symbols that occur nowhere in the input marking the self-commentary parts, omitting the prefix you used to get gpt3 to output it, and train a new network (GPT3') on the augmented data.
Then you could make an interface that uses GPT3' and hides the self-commentary from the users. As GPT3 writes it will have a persistent memory that can last as long as the document goes on, updated by itself. Effectively it gives it sparse access to the entire history, but the network itself controls the shape of the access.
[Plus a nice thing about GPT generated text is that you can store its confidence too, and use that to weigh the training so that you penalize it less for mispredicting stuff it was unsure of.]
You wouldn't have to do anything special to teach it to write this commentary because we already write commentary in English and GPT3 already knows how to do it.
Maybe a little more engineering would be useful to guarantee that it will write commentary blocks often enough, but it could be as simple as making sampling prefer emitting an internal monologue block with an increasing bias as the last one approaches falling out of the window.