An actually large context window is impossible due to how LLM attention works under the hood.
This has obvious issues since you're now losing information from the now unseen tokens which becomes significant if your context window is small in comparision of the answer/question you're looking at. That's why companies try to give stupidly large context windows. The problem is they're not training on the large context window, they're training on something smaller (2048 and above). Due to how attention is setup, you can train on a small amount of context and extrapolate it to any number of tokens possible since they train via ROPE which trains the model because on words and their offset to the neighboring words. This allows us to effectively x2,x3,x10,x100 the amount of tokens we generate vs train with with some form consistency BUT still cause a lot of issues consistency wise since the model approaches more of a "this was trained on snippets but not the entire thing" situation where it has a notion of the context but not fundamentally the entire combined context
The only real mistakes it makes are some model specific quirks, like occasionally stripping out certain array index operators. Other than that, it works fine with 150.000 token size conversations. I've gone up to 500.000 with no real issues besides a bit of a slowdown. It's also great for log analysis, which I have maximized to 900.000 tokens.
The limiting factors are typically: 1. Often there are latency/throughput requirements for model serving which become challenging to fulfill at a certain context length. 2. The model has to be _trained_ to use the desired context length, and training becomes prohibitively expensive at larger contexts.
(2) is even a big enough problem that some popular open source models that claim to support large context lengths in fact are trained on smaller ones and use "context length extension" hacks like YaRN to trick the model into working on longer contexts at inference time.
And sure maybe not 2mil of it is usable, but they're reliably pushing the frontier here.
For example when querying a model to refactor a piece of code - would that really work if it forgets about one part of the code while it refactors another part?
I concatenate a lot of code files into a single prompt multiple times a day and ask LLMs to refactor them, implement features or review the code.
So far, I never had the impression that filling the context window with a lot of code causes problems.
I also use very long lists of instructions on code style on top of my prompts. And the LLMs seem to be able to follow all of them just fine.
https://wandb.ai/byyoung3/ruler_eval/reports/How-to-evaluate...
>Gpt-5-mini records 0.87 overall judge accuracy at 4k [context] and falls to 0.59 at 128k.
And Llama 4 Scout claimed a 10 million token context window but in practice its performance on query tasks drops below 20% accuracy by 32k tokens.
Here is an experiment:
https://www.gnod.com/search/#q=%23%20Calcuate%20the%20below%...
The correct answer:
Correct: 20,192,642.460942328
Here is what I got from different models on the first try: ChatGPT: 20,384,918.24
Perplexity: 20,000,000
Google: 25,167,098.4
Mistral: 200,000,000
Grok: Timed out after 300s of thinkingYou wouldn't ask a human to do that, why would you ask an LLM to? I guess it's a way to test them, but it feels like the world record for backwards running: interesting, maybe, but not a good way to measure, like, anything about the individual involved.
Then there's the question of why not just build the calculator tool into the model?
Tested this on the new hidden model of ChatGPT called Polaris Alpha: Answer: $20,192,642.460942336$
Current gpt-5 medium reasoning says: After confirming my calculations, the final product (P) should be (20,192,642.460942336)
Claude Sonnet 4.5 says: “29,596,175.95 or roughly 29.6 million”
Claude haiku 4.5 says: ≈20,185,903
GLM 4.6 says: 20,171,523.725593136
I’m going to try out Grok 4 fast on some coding tasks at this point to see if it can create functions properly. Design help is still best on GPT-5 at this exact moment.