If you use open weight models, you can run several agents for the same budget. From what I've seen of Claude, there is no difference for most daily work. Claude makes many mistakes too. The real thing that sets open weights appart is that you can leverage multiple model families from the same token vendor. Using multiple families has outsized impact imo. Like... don't have Claude review the code it wrote.
I time slice between them, have them do a lot more reviews of each other, and spend more effort on thinking about design and checking their code.
I'd argue doing audits and having agents clean up their own slop, then creating instructions / guidelines from those, will reduce the problem long-term. They very much will mirror the surrounding code the code is consistent itself and the core points are called out in a little AGENTS.md