I tend to define it "better at solving, worse at assisting phenomenon". Which doesn't properly show on benchmarks that only focus on the solving part.
Though my feeling, no proof, is that the opus/fable today is not what it was months ago. there was a time for about a month where opus was incredible. Just incredible but as fable started to move out i swear to god it feels like sonnet now. Fable feels like opus used to but costs more.
Stick the "Never suppress errors" section into your Claude.md, this will never happen again (works for me with Python/Flask, ymmv for other languages).
I don't like OpenAI as a company, but they appear to have QA, and that is probably enough to get me to switch.
Basic stuff about features that are more than a week old just get no attention at all. From the outside Athropic seems to be a clear feature factory.
IMO that's exactly why it's a bit better at actual problem solving.
You absolutely do not "always have to correct" Codex. I'm not sure what you're doing, but I'd say 80-90% of its edits on my side it doesn't need any revisions.
this has been my experience with Codex as well, and I have to fix its mistakes every single time. But recently, I literally threw away three hours of work because it kept adding hundreds of lines to my code base. When I restarted the entire work using Fable and Opus, it was like night and day.
Obligatory YMMV, maybe your prompting style fits gpt better. We forget that this matters a lot