The top models, when used in a harness, will likely catch those errors or somehow manage to get to the right answer at some point, but those tests test mostly how likely a model is, when used via API, to give the correct response.
Opus 5.5 is not trash, still in the top models, but Anthropic models always have struggled with instructions following and refusing to answer questions. That being said, most tests are basically questions or simple tasks, and models could either do them ok or not. Nowadays the models and how good they are in practice is given more by the harness, than the model itself. I do think I should probably find a way to test the models including their harness, and to do so for a complex, long-running task, the so called "agentic" use-case.
Also, Gemini models have the best all-around knowledge, they are way above other models in general knowledge and domain-specific knowledge. No tests have web search enabled, and most modern models are indeed optimized for that use-case nowadays.
tl;dr: the leaderboard simply shows, given any simple question or programming task, which model is most likely to get the answer right.