Seems intentionally misleading.
Gemini Ultra scored 90% which is better than GPT-4.
This reads like a paid-for press release from Microsoft to pretend like they're almighty and Google is incompetent.
Edit: BTW, more Mistral benchmarks here: https://docs.mistral.ai/platform/endpoints/ TIL Mistral Small outperforms Mixtral 8x7B.
As it stands best LLM available by API by Google is far behind GPT4.
Most of my usecases are logic based on embedded content in the prompt and nothing available to me beats GPT-4 there.
> generally available through an API (next to GPT-4)
edit: not sure why I am being downvoted. I am 100% sure the way they structured it was meant to say "we are doing great, but not as great as openAI's work, which we are not trying to compete against". I guarantee there were discussions on how to make it look as to not appear that way.