DL practitioner for a decade here:
The OP doesn’t explain anything. It just vaguely talks about a few things that might break when scaling context. But that means nothing.
Take for example sinusoidal embeddings they talk about. Of course it breaks for large contexts but in no one uses it. GPT uses learned positional embeddings so the entire section is irrelevant.
Copy this for pretty much everything else.
Being an expert in a field has never been this exhausting