I think the key is to give them a nice assortment of self-verification tools, an AGENTS.md or reference document that they're encouraged to routinely check, and asking the planner to be thorough with the ACs but give the model some space.
The planner routinely finds issues with the worker's output, but that's what it is for.
All the more reason to favor local models under your control, as once you find that sweet spot model, no one can change it, upgrade it, align it, take it down or otherwise harm the time investment you made it making it your own.
I can't really believe no one understands, after decades, how valueable a rock solid development environment is.
I would like to see some development where proof of authenticity certs are generated alongside the actual output of the model. Prove to me (or at least claim to me liable to breach of contract) that this output was generated by FP8 DeepSeek V4 Pro 0813. Not some cheaper quantization of the model.
For this however, a comparatively much simpler task, tarra-high works fine.
DeepSeek is okay for random API-based stuff, as it's cheap.
Local open models running on a 5090 are hit or miss. I feel that most GGUFs/quants are awful...
I wonder if it is because of watermarking.
That was true before they announced the watermarking, I'd already started to back off of using Opus as much because I like to understand what the model is doing and have it write documentation I can use to reproduce its results, but maybe watermarking was already in there unannounced.
And it retypes it for a human, I'm doing this more and more lately.
I think it could be the watermarking, but at this point they might be deliberately complicating the prose so that we ask clarifying questions and that leads to more token spend.
Install the Superpowers plugin.
Behold.
I layer on CodeRabbit for PR Review and it’s just so solid.
This week I removed it because it now gets in the way of the frontier models.
The harness is everything. If I just threw something like Codex at this and said "good luck" I wouldn't make it beyond 5-10 interactions. I tried that already. Carefully designing the views and tools over the environment is where you can go from 50% to 99.9999%.