Guide coding is more forgiving because the diffs are small enough that you can just try something and veto it if it's not good. Qwen 3.6 a3b makes more mistakes than DeepSeek but it's free. I'd imagine the next ~32gb-vram-class MoE model from Qwen will close the gap.
The real deciding factor for me is inference speed.
I use VSCode insiders with their BYO model configuration. I stage little bits at a time in a tight prompt-review-prompt-review workflow. I occasionally use dedicated harnesses (DeepSeek harness, Codex, etc) because they have better tools and outcomes than VSCode's built-in harness for longer horizon tasks - though I lack the ability to highlight a block of text and say "add error handling" or similar.
When building visual applications, desktop harnesses are better because it's a bit easier to send screenshots to the agent.
I love Zed editor but its AI review features are lacking compared to VSCode. I go back to it frequently and am ready to switch over when the team resolves the usability issues.