Impressive how well Grok performs in these tests. Grok feels 'underrated' in terms of how much other models (gemini, llama, etc) are in the news.
less people using them.
I think the actually-relevant issue here is that until last month there wasn't API access for Grok 3, so no one could test or benchmark it, and you couldn't integrate it into tools that you might want to use it with. They only allowed Grok 2 in their API, and Grok 2 was a pretty bad model.
Also, only one out of the ten models benchmarked have open weights, so I'm not sure what GP is arguing for.
not talking about TFA or benchmarks but the news coverage/user sentiment ...
Gemini frequently avoids discussing health problems, which likely hurt its scores. My guess is any censorship was considered a fail.