Anthropic made it that way, and I'd say the lower score is accurate.
Anthropic made it that way, and I'd say the lower score is accurate.
I just think that this benchmark measures what it say it does, and if the model is unable or unwilling to deliver, it's reflected in the score.
(incidentally I found that generating code with Sonnet agent and having Opus orchestrate and manage the process works best for me - right model for the right task - the results run in the places and the way it's okay. No one outside uses it :p)
I agree with your point - IMO lower Opus score in these suggests that in general it's worse for these tasks. Not that it's a worse model in general.
But yeah, it does, but from my perspective it measured the Opus performance - subpar in some tasks because it downgraded itself rerouting to a much weaker model.
So both you're right, it matters because it wasn't the model examined, and it doesn't matter because the score reflects nerfed experience resulting in nerfed results.