A common tactic is to used a big brain model like Opus for planning and reviewing, and a cheaper model for execution.
A common tactic is to used a big brain model like Opus for planning and reviewing, and a cheaper model for execution.
Other than that, I think they are cheaper last time I compared.
Otherwise if you don't care about your data being used in training you might as well use the best, which is the US models on a subscription plan. If you do care, use downloadable weight models served by reputable providers, or selfhost.
I've experienced this firsthand and now I generally pin to providers that I trust on openrouter, or just pony up and pay for the real thing.
In my experience (and I've been trying this a bunch): smart planner + dumb executor produces worse code with higher spend than simply using the smart planner to do both.
It's easy to understand why:
- If the planner has truly thought the issue through, properly designed the solution, solved all of the emergent problems, then the final "write" of the code is just a few more output tokens.
- If the planner has NOT truly planned the issue completely, then you're letting a substantially dumber and less capable model make significant decisions, and trusting its problem solving, without having a better model check it.
If you're highly cost conscious (paying for your own tokens and not making any money) then you have no choice but to trade your time and effort for tricks like this to save money by lowering the quality of your output.
But if your employer is paying for tokens: just use the smarter model. You save your time preventing re-work and reducing code review, you save your employer money (primarily from the cost of your own labor and reduced rework), and you get a better output every time (Opus 5.5 mogs Deepseek 4.1 flash in every single way except cost).
Etymology: probably from AMOG Alpha Male of the Group
First seen: 2018
I also agree that its a big mistake to have a flash model implement without a strong model reviewing.
I have Opus plan, Deepseek implement the code, and then review with Opus [1]. In this workflow I am saving a lot of money by having Deepseek do the implementation. Note that the review back-and-forth is fully automated [2], so it doesn't take any extra attention from me.
[1] https://github.com/gregwebs/skills-sdlc/tree/main/skills/implement
[2] https://github.com/gregwebs/skills-sdlc/blob/main/skills/code-review-with-followup/SKILL.mdI have tried your workflow many times, and simply letting Opus do the implementation costs much less than wasting hundreds of millions of tokens letting deepseek and opus go back and forth and back and forth. And bonus, my project finishes in 5 minutes instead of 20.
No, the only open weight model that really makes sense for me is Qwen3.8-Flash-Next, but it is mainly because I can run it locally with reasonable speed (prefill between 650-1400t/s generation between 22-50t/s depending on number of slots/users I configure).
This is the first model that truly competes with Opus 4.8. I'd say it may be better than Opus 4.6 on programming.
But it is very verbose when it comes to reasoning tokens. The more difficult the task the more verbose it is. Certain very hard tasks that take opus 4.8 400k tokens take Qwen3.8-Flash-Next 2M tokens... But it finishes them.
And what you loose on the generation speed you get back on input caching you can keep on for weeks.
It really depends on the workload.
I do have an /implement-simple workflow to skip the planning phase, but even that doesn't skip the review.
Are you doing your own intensive reviews of the model code? Can you share the prompts you are using as I have?
My bar for what models produce without human intervention is much lower defect than what a human would produce. The human interaction is mostly to guide the design and then the review burden is very low. I suspect your bar for what agents produce is lower- you are taking more of the review burden. I also suspect that you are measuring time more than actual cost since your employer is paying and that you are comparing to Sonnet rather than DeepSeek (DeepSeek 4.1 again is 20-40x cheaper than Sonnet). You mention hundreds of millions of tokens (my reviews don't use that much), but even that costs ~$1 on the DeepSeek side.
I think you are taking exactly the right approach at your employer given the cost is free and you only have access to Anthropic models.
One thing that I have found is that as the frontier models get better there is less need for agents with specialized personas. I actually don't don't use those anymore- I just use agents that have different models and reasoning levels. I have a generated CODING_STANDARDS.md document and a skill for architecture design and a skill for implementing testing [2] that are referenced by a single reviewer. I do implement a 2-pass review though [3].
I would be interested to know if you have found anything similar as models get better. It seems though that you are sharing a single exploration and then sharing the context across the specialized reviewers to dramatically reduce the cost of your approach. Does this have to be in the harness- that is if you write out the shared context to a file does that increase your costs a lot?
I also wonder how intensively are the models able to test their changes? The number one quality improvement I have found is not review but having the model properly test its code. I have a skill that is helping [4], but I also have to spend time to establish a pattern of testing with tools beyond just unit tests. The testing takes significant effort, and this is again where the cost savings of DeepSeek shine.
[1] https://github.com/mattpocock/skills/blob/main/skills/engineering/codebase-design/SKILL.md
[2] https://github.com/gregwebs/skills-sdlc/blob/main/skills/verify/SKILL.md
[3] https://github.com/mattpocock/skills/blob/main/skills/engineering/code-review/SKILL.md
[4] https://github.com/gregwebs/skills-sdlc/blob/main/skills/verify/SKILL.mdAlso if you aren't hitting capacity.
Off-work, I use LLMs regularly for both design/coding and non-technical work, but the volume is not enough to trip the weekly limits, and rarely enough to trip the daily limits. So I just go with whatever's current best SOTA available on my Claude & ChatGPT subscriptions and don't worry about limits. If I hit one, I do some household stuff or relax for a few hours (or just turn in for the day), and then the limit is refreshed.
For me, what usually creates the high cost of implementation with a larger model is the validation process, not the writing of code.
I would get excited to see it finish writing code with so little usage, but then it would gobble up ten times as many tokens on validation.
I've tried to instruct it to keep the validation light with the intention of doing batches of deep validation after a few tasks, but it couldn't stop itself from doing heavy validation on each task no matter how I rephrased the instructions.