Any examples?
And GPT 5.6 Sol over engineers just about everything. No LLM is perfect, its about learning the issues with each LLM and figuring out if you can live with it. Knowledge means that you can anticipate if it tries to pull something funny, and harness it against that behavior.
It's horrible advice given what we've seen consistent: changing alignments, changing guardrails, changing system prompts, changing inference priorities, etc.
Anyone who relies on these for their work product is chaining themselves to a matrix multiple of indetermintism.
But yes, I also think it's not the greatest model for programming. On the other hand, for agentic tasks that are not programming related it's hard to beat Opus 4.8. It can try different things and pivot even when the user is not great with prompting. 5.0 seems to not be worse, but definitely wastes more tokens and costs more.
I'm doing data science stuff so it isn't super complicated code; it is about applying valid statistical procedures and techniques. Still, on the code part, Opus 5 had a lot of trouble merging 2 branches yesterday...
On a tangent, I am beginning to understand why we have replication crisis in academia. I thought C++ was full of footguns; it has nothing on statistics. With statistics, you don't get a compiler error or a crash when you hold it wrong.
I've read that Opus 5 has more success if used with Fable as orchestrator, but since I refuse to pay for subscription access to Fable, I'm now trying Opus 4.6 as orchestrator (also since that was the version that got me loving Claude), and Opus 5 low effort as implementor.
Sometimes it will also not careif a unit test fails, calling it "errant" or "legacy" or something, where it then removes it from the list (not sure how to get around that one, other than a read only launcher).