In practice, the 1M context of Gemini 2.5 isn't that much of a differentiator because larger context has diminishing returns on adherence to later tokens.
A good deep dive on the context scaling topic in general https://youtu.be/NHMJ9mqKeMQ
Sure, gpt-4o has a context window of 128k, but it loses a lot from the beginning/middle.
I think it also crushes most of the benchmarks for long context performance. I believe on MRCR (multi round coreference resolution) it beats pretty much any other model's performance at 128k at 1M tokens (o3 may have changed this).
At 500k+ I will define a task and it will suddenly panic and go back to a previous task that we just fully completed.
Particularly for indie projects, you can essentially dump the entire code into it and with pro reasoning model, it's all handled pretty well.
That said I have noticed that if I try to give it additional threads to compare and contrast once it hits around the 300-500k tokens it starts to hallucinate more and forget things more.
Other tools may drop some prior context, or use RAG to help but they don't force you to start a new chat without warning.
(same as sonnet 3.7 with the beta header)