My experience is Codex is much better but less creative. I use its agent through GitHub Copilot or the agent interface from JetBrains. Try the GPT Sol 5.6 on medium, it gives me good results
I think the tradeoff to Claude being so needy is that if you let models just run away with an inaccurate or incomplete understanding of what to do, they can go really far off the rails AND spend a lot of time/money doing it AND come back with something that literally doesn’t make sense or doesn’t work.
I prefer dealing with Claude’s reliable cringe to the aloof model that tries to play it cool when it needs help.