If GPT 5 truly has 400k context, that might be all it needs to meaningfully surpass Opus.
To get great results, it's still very important to manage context well. It doesn't matter if the model allows a very large context window, you can't just throw in the kitchen sink and expect good results
But is it really 272k even if the output was say 10k? Cause it does say “max output” in the docs, so I wonder
I found 100k was barely enough for a single project without spillover, so 4x allows for linking more adjacent codebases for large scale analysis.