I haven't tested the new version, and I still have a bunch of tokens on my token plan, so I might give 2.6 a go in some similar tests to see if it still has the analysis paralysis problem of 2.5.
The non-pro version is mostly useless
The bigger issue is that the use cases and harnesses for models is infinite, which is hard to compress into benchmark numbers that actually apply to you.
Everyone is benchmaxxing, desperate to sell, and almost nobody except the labs is doing actual science on the results, so harnesses tend to be chosen on voodoo and hunches, like which company made it. There isn't necessarily a good alternative though, bearing the cost of being a harness researcher is probably not many people's goal.
> The model matters more than the harness anyway
>
> Everyone is benchmaxxing
>
> ...harnesses tend to be chosen on voodoo and hunches...
I get what you're saying, but their graphic on performance here uses the exact same model with different harnesses and definitively shows that there is a significant difference in both accuracy and cost. The whole point of their technical implementation and design decision here is to highlight that it's not "voodoo and hunches", but observable data.Fable 5 on Claude Code scored 61.8% at a cost of $248.05 while Fable 5 on OpenCode beat it at 66.3% at $73.42. The same model, the same benchmark; only the harness is different with a ~5 point difference in accuracy while costing significantly less. So if we are to believe the author and these results are repeatable, then it would seem that the harness matters.
The point of this framing here is specifically to address 1) benchmaxxing by using the same model, 2) NOT choose a harness on "voodoo and hunches" by using actual data to back the assertions. Your comment feels misguided and completely hand waves the actual data points here.
Cost efficiency is a plus
There's no canonical Pi harness, it's too variable. There might be a world where something Pi based is the target, like OMP, but then you have to hesitate when you start extending the harness because you don't know what effect any given extension will have on model performance.
Their excuse was that it was a file indexing optimisation and they forgot to create a toggle or setting around it. As an apology they open sourced the harness, but wiped out the history.
It sounds to me like a good distillation technique if the project has hundreds of commits written by Claude and detailed commit descriptions. Also for completeness Grok did the exact same thing two months ago.
https://www.theregister.com/security/2026/09/22/zai-says-sor...