Half the cost per task compared to Opus 5, comparing high effort to high effort. That's just really nice.
Edit: https://artificialanalysis.ai/models/claude-opus-5-5?models=...
Edit: https://artificialanalysis.ai/models/claude-opus-5-5?models=...
It's also less clear what a lot of their metrics mean. Does Cost per Task include only things that can be verified to work and passed? As best I can tell, it does not.
I'm less concerned if one model's cost per task is $0.10 and another model's cost is $1.50 if the $0.10 task got it right 1% of the time and the $1.50 model got it right 66% of the time.
An equalized / weighted cost/time per task is much more valuable - being massively penalized for taking a lot of time and ultimately not passing when OTHER models did pass.
Not sure about how adaptive reasoning works though as they mention adaptive reasoning for every reasoning level