"Opus 5 is a really bad model"
twitter.com
twitter.com
Agent to agent communication does appear to have improved considerably. I run the loops to the full context window and with proper working docs I don’t even notice a difference after compaction.
In my case, I have a semi-autonomous loop where Claude writes some code, and uses `codex exec` to do an adversarial review. I had what should have been a trivial feature go for 13 rounds of review/fix before I stopped it, where each round was just flip-flopping the same logic back and forth to try to make the tests pass. Codex kept (correctly) re-identifying the issues Claude was flip-flopping on. Codex even suggested fixes that would have worked; Claude ignored them repeatedly. I never saw anything remotely this bad on Opus 4.8.
Additionally, I have a CLAUDE.md instruction to not silently defer anything, and ask me anytime it wants to do so. Opus 4.8 paid attention to this rule the overwhelming majority of the time. Opus 5 seemingly cannot be bothered.
I have tried updating my CLAUDE.md according to Anthropic’s recommendations for Opus 5, but it doesn’t seem to have made any difference.
<system-reminder> IMPORTANT: this context may or may not be relevant to your tasks. You should not respond to this context unless it is highly relevant to your task. </system-reminder>
It's not like users can prove anything about a remote black box anyway.
I really miss the 4.5 era, that was a magic time.
It seems like every time I see a screenshot of someone’s session it’s full of “u fix yet” or “why didnt you fix it” prompts. I thought I was only seeing posts from those people because those people are the ones that get their repo wiped by AI, but at this point it seems I’m the last of a dying breed or something. I even saw an OpenAI model launch that used ‘drunk text message’ formatting in the examples.
Edit: someone downvoted this comment, which I can only guess is because they’re used to people that use “honest question” as a prefix for a rhetorical or insulting statement dressed up like a question. That’s not what I’m doing. I don’t have the capacity to care how other people interact with their tools, it doesn’t affect me at all. I’m truly asking because I see it a lot lately, but not in my circle, so perhaps I’m a generation removed from it. My theory is that I learned to type on a keyboard, where autocorrect and capitalization was performed manually, but later generations learned primarily on a phone that handled all that for them. Perhaps switching to a keyboard to use Claude didn’t naturally engage the need to capitalize letters and such.