I'm surprised that Opus 4.5 is better than Opus 4.6 and Sonnet 4.6 is even better than Opus 4.5 (and 4.6). Shouldn't Opus 4.6 be the best of the Claude models?
I can’t really tell the difference between the two models for the things I do any more.
That's a nice benchmark + website and wow ChatGPT scores worse than I thought.
That explains why I intrinsically "trust" Sonnet 4.6 the most.