There's no way GPT 5.5 is better than GPT 5.6 Sol, and gemini-3.6-flash is better than both of those. I wouldn't trust this benchmark at all.
maybe if you mentioned 3.7-flash it might have been slightly more believable (but still false).
The reason for both of those things is, as you point out, the benchmark is very obviously saturated