that's interesting... i've been noticing similar issues with long context windows & forgetting. are you seeing that the model drifts more towards the beginning of the context or is it seemingly random?
i've also been experimenting with different chunking strategies to see if that helps maintain coherence over larger contexts. it's a tricky problem.