Your mental model of benchmark scores is off.
Some tasks within the benchmark are much easier than others. The hardest several tasks often have vastly different difficulty levels. Often, the hardest few tasks are literally impossible; malformed problems due to poor curation, often.
Imagine you've got a basketball robot, and one way you test it is on the Three Pointer benchmark. It tests the robot's ability to shoot a three pointer from 20 feet, 25 feet, 30 feet, 40 feet, 50 feet, 60 fee, 75 feet, 100 feet, 200 feet, and 182 miles.
Is a robot that scores 90% on this benchmark 90% as capable as one that scores 100%?