That’s why it’s probably a good idea to never stick to one model but kick off a fleet on the same tasks and in parallel and drive consensus.
At least, that’s what I’ve found to be useful by pitting claude/codex/etc against each other to keep them a bit more honest.