I get drastically different tool call failure rates using Claude SDK vs OpenCode using Qwen 3.6 models
I get drastically different tool call failure rates using Claude SDK vs OpenCode using Qwen 3.6 models
This feels like a horrible failing of the models to generalize, then - both basic and intermediate tasks should be possible to do with Claude Code, OpenCode, Pi, ZCode, Kimi Code, Dirac and tbh any other mainstream or even slightly niche harness. Not doubting the claim itself, there's a reason why good benchmarks include the harness.
For coding specifically, I'm not sure this is still true. Given the heavy use of RL to improve coding performance, I'd expect the harness to be important as it defines what tools the model is rewarded for using.
It's crazy how over the past years a field originating from math ends up succumbing to subjective feels.
You can use effectively any harness and get good results. Harnesses are mostly placebo.
To put that in perspective, the difference between GPT-5.6 Sol Max and 5.6 Luna Max is 8 points. That's a lot of extra performance that you can get for free just by using the best harness.