I wonder why they didn't test Gemini 3.5 Flash (High).
Out of curiosity how are benchmark runs generally funded? It would obviously be great to test them all on all reasoning levels and in and out of their native harnesses. Maybe even in pi / opencode / cursor but I get this would get prohibitively expensive unless you have funding or free tokens.
Thanks for your efforts thus far. Looking forward to seeing more.