How are you using GPT-2 with an expanded context window? I was under the impression that the maximum context window was fixed.
Specifically, the code checks whether the model's shape is greater than the shape from the snapshot on disk. If so, it repeats the shape from the snapshot on disk N times to fill the expected greater shape.
At that point, you can just set context window to a larger value, then train.
Generating more tokens seems to work up to a point -- you can probably generate up to 1050 tokens with this technique. But at a certain point, more tokens = gibberish.
The cure is to train the new wpe layer the same way you'd train the smaller one. But this also means you don't have to start training from scratch.