Sometimes I can't really tell.
Input $0.28 / 1M tokens cache miss Output $0.42 / 1M tokens
Via synthetic (which otherwise looks cool):
Input $0.56/mtok Output $1.68/mtok
So 2-3 better value through https://platform.deepseek.com
(Granted Synthetic gives you way more models to choose from, including ones that don't parrot CPC/PLA propaganda and censor)
unfortunately it doesn't support local models but they're too slow for coding anyway.
GLM is maybe slightly weaker on average but on the other hand it's also solved problems where both CC and Codex got stuck in endless failure loops so for the price it's nice to have in my back pocket. I also see some tool use failures sometimes that it always works around which I'm guessing are due to slight differences with Claude.
It's ok for documentation or small tasks, but consistently fails at tasks that both Claude or Codex succeed at.
Compared to the anthropic offering is night and day. Claude gets on with the job and makes me way more productive.
Which model were you using? In my experience Gemini 2.5 Pro is just as good as Claude Sonnet 4 and 4.5. It's literally what I use as a fallback to wrap something up if I hit the 5 hour limit on Claude and want to just push past some incomplete work.
I'm just going to throw this out there. I get good results from a truly trash model like gpt-oss-20b (quantized at 4bits). The reason I can literally use this model is because I know my shit and have spent time learning how much instruction each model I use needs.
Would be curious what you're actually having issues with if you're willing to share.
Is just strange to me that my experience seems to be a polar opposite of yours.
I can one-shot new webapps in Claude and Codex and can't in Gemini Pro.