To me, the play is: open weight on a provider like BaseTen (solid performance, low price point), or pay up for Gemini3.1 Pro if you need it.
But at their high price and low-ish quality, OpenAI models just aren't in the conversation right now without heavy incentives, e.g. via Azure.
Crazy, TBH. Curious if others find the same thing?
My thinking is:
codex - best harness
opencode - best ux/dx
claude - best value for moneySo I would agree with you, it is not great.
So yes, while we can work with any of these models to get them to do what we need eventually -- e.g. with prompt tuning to their particular style, adding more examples, or breaking tasks into smaller steps, etc. -- their instruction following has a huge impact on how quickly we can move as a team.
When I say "stinks", for me, if we do three rounds of optimization and testing and a model is still performing inconsistently across a class of related traps then using that model is going to slow us down, and I think it stinks.
In my experience, gemini3.1pro tends to work very consistently with light nudging, GLM with 2-ish rounds of optimization, and for GPT5.4, well it provided no improvement over prior models and would slow us down over the others meaningfully ... and costs too much for the effort.
So, meh, so I still think it stinks, skill level considered.
We still have to see what Anthropic has cooked though
Anthropic models are still the best for me -- as long as you don't ask them to do something they don't want to -- but also, way too expensive for bulk pipeline processing. So I keep it to coding and Coworking...