No benchmark comparisons to other models, especially Gemini 2.5 Pro, is telling.
They are reporting that GPT-4.1 gets 55%.
Andrej Karpathy famously quipped that he only trusts two LLM evals: Chatbot Arena (which has humans blindly compare and score responses), and the r/LocalLLaMA comment section.
In practice you have to evaluate the models yourself for any non-trivial task.
This is pretty common across industries. The leader doesn’t compare themselves to the competition.
[1] https://blog.google/technology/google-deepmind/gemini-model-...
[2] https://ai.meta.com/blog/llama-4-multimodal-intelligence/