You can see this when you use Anthropic Claude which has a 100K context length today.
I have a feeling this will become a non-issue in the near future as the models are further trained with this in mind.
Take an undertrained model for example: It starts becoming incoherent as you approach the context length - I have a theory that OpenAI models have been running at a larger block-size than presented for a while now - for example, "4K" models actually had 8K context but capped at 4K as anything beyond starts becoming incoherent: Reason being, you train to around 5K and don't let the user go near that section of the model and it gives the impression that the entire context block is 100% functional.
The solution is trivial: You bootstrap the models by having them generate training data after they reach a certain point.
I wrote one from the ground up (PyTorch only) with the intention of having it perform in constrained environments and these have been my findings over the last few months.
Yes. When you say something stupid, ChatGPT won't forget it as easily...