Everyone who survived 7 rounds of multi model reviews and they still keep finding mediums in their PRs is not in the least surprised. These things are not oracles - they miss stuff all the time even when told to look.
well not token usage, but revenue. their costs for this work would have been astronomical in their own service tier because i bet the context was way larger than anything they even offer.
tweaking context size is the main, or only, "strategy" they have for cost/revenue. and is the reason new trained versions continue to generate hype: you need data in training because you cannot have it in context