Your typical LLM benchmarks simply do not test or use large context sizes.
We need benchmarks for tasks that requiere large context sizes (like recalling facts and understanding them in long stories). I'm sure OpenAI have internal benchmarks for these tasks.