In one instance, I asked it to optimize a roughly 80 line C# method that matches some object positions by object ID and delta encodes their positions from the previous frame. It seemed to be confused about how all this should work and output completely wrong code. It has all the context it needs in the file and the method is fairly self-contained. Other models did much better. GPT-5 understood what to do immediately.
I tried a few other tasks/questions that also had underwhelming results. Now I've switched to using GPT-5.
If you have a quick prompt you'd like me to try, I can share the results.
But they definitely don't taking into account whatever prompts the tools are really using (or ms is using a neutered version to reduce cost). So I would agree with the suggestion. Using sonnet through copilot seems very very different than cursor or cline or Claude code.
Using the same exact model, Copilot consistently often fails to finish tasks or makes a mess. It is consistent at this across ides (ie using the jetbrains plugin generates nearly identical bad results as vscode copilot). I then discard all it did and try the exact same (user) prompt in cursor or Claude code or cline with the same model and it does the same task perfectly.
Perhaps it shouldn't be surprising; after all, we do want the LLMs to listen to the prompts and act differently. And, the Claude team will presumably be tuning both Claude and Claude Code's prompts to each other optimize their own experience, so it's perhaps not surprising that Claude + Claude Code's prompts well together.
However, if I'm not detailed with it, it does seem to make weird choices that end up being unmaintainable. It's like it has poor creative instincts but is really good at following the directions you give it.
I spend a lot of time planning tasks, generating various documents per pr (requirements, questions, todo), having AI poke my ideas (business/product/ux/code-wise) etc.
After 45 minutes of back and forth in general we end up with a detailed plan.
This has also many benefits: - writing tests becomes very simple (unit, integration, E2Es) - writing documentation becomes very simple - writing meaningful PRs becomes very simple
It is quite boring though, not gonna lie. But that's a price I have accepted for quality.
Also, clearing the ideas so much before hand often leads me to come with creative ideas later in the day, when I go for walks and review mentally what we've done/how.
People tend to hate Claude Code because it's not vibe coding anymore but it was never really meant to be.
And if it's gets confused, needs clarification, or has its own initative - I want it to stop and ask.
Oh and it needs to be fast it's tokens per minute should be as fast as I can read what it generates (and I can read boilerplate-y code quite fast), and it shouldn't stop and think on every prompt, only when it needs to, and it should be much faster and granular in backtracking.
The loop of waiting on the AI then having to fix and steer it constantly as it doggedly follows its own ideas has really taken the enjoyment out of vibe coding for me.
Also keep in mind that many employees are not paying out of pocket for LLM use at work. A $1,000 monthly bill for LLM usage is high for an individual but not so much for a company that employees engineers.
They're impressive despite that. But if Sonnet is $20/month and I have to intervene every 3 minutes, while Opus is $100/month and I have to intervene every 5 minutes? ¯\_(ツ)_/¯
Inverting the problem, one might ask how best to spend (say) $5,000 monthly on coding agents. I don't know the answer to that.
So do engineers.
The difference is that IRL engineers know a lot about the context of the business, features, product, ux, stakeholders, expectations, etc, etc which means that the hand-holding is a long running process.
LLMs need all of these things to be clearly written down and specified in one shot.
I think this is the whole reason not to compare it to Opus...
I believe Opus starts at $20 a month, similar to GPT5 if you want more than just cursory usage.
Or am I missing something?
Claude Opus 4.1
Most intelligent model for complex tasks
Input $15 / MTok
Output $75 / MTok
Prompt caching
Write $18.75 / MTok
Read $1.50 / MTokIt would be useful to be able to easily compare what it costs across the big providers: Gemini, Grok, Claude, ChatGPT.
If you want to use Opus in claude code, you've got to get the $100/month plan - or pay API prices. And agentic coding uses a lot of tokens.
> Our smartest, fastest, and most useful model yet
I'd say it's definitely supposed to be the best, it just doesn't deliver.
> I'd say it's definitely supposed to be the best, it just doesn't deliver.
What part of "Our" is difficult to understand in that statement? Or are you claiming that OpenAI owns another model that is clearly better than GPT-5?
I would suggest reading the entire comment thread before attacking people.