Sol 6.1 scores one point less than Gemini 4 on intelligence AND costs less than half ($0.72 vs $1.99) per task.
Additionally, if you are using OpenAI you have the option to pay a bit more and get Astra which - despite the benchmarks - does outperform Sol on some things.
Also, people are - rightly - very wary of Google's benchmaxxing tendencies. I think lots of people remember Gemini 3.0 (I think?) which benchmarked amazingly, but as soon as you used it would go off-track and needed constant babysitting if you wanted to use it for agentic work.
Via the API. The $200 OpenAI plan just got cut and most say general quotas got cut before that so for users on a plan the numbers might be different.
https://www.bridgebench.ai/nerf-bench
https://marginlab.ai/trackers/claude-code/ https://marginlab.ai/trackers/codex/
Methodology section in some of these benchmarks doesn’t say if they use subscription or API. API usage may not be nerfed as much as subscription.
They must be using the same lever to “pace the frontier”. All of the best effort models from different companies have similar scores. There is no standard definition of “max” effort level.
the deal has a 30-day cancellation policy, and they raised a bunch of debt around the same time to fund their own datacenter expansion.
Anthropic has some of the lowest usage per $ in general, not sure what you're taking about.
Opus and Sol usage levels vs the API are currently roughly the same, but Opus 5.5 outperforms at low and medium effort levels.
It's all relative. Some people just can't use up their quotas with their normal usage.
If it's smart enough to do the job then it won't matter that Opus is smarter. At the right price and performance, at least.
Honestly, I think this will be a trend in 2027 when all those models become "good enough" for elite coding and whatever. I predict they'll have to branch out more and build their platform to differentiate themselves from each other, maybe even in terms of branding and trust, marketing towards various demographies, youth vs elderly, students vs employees, etc.
I pay for lots of models because they’re good at different things: $20 a month each for Grok and GLM have easily paid for themselves by finding bugs that my main work models didn’t, but I’m yet to have any Gemini model find a real bug, and Gemini’s results for general work will sometimes border malicious compliance, when it’s not having a hissy fit about some imagined issue.