The benchmark would suggest that but if you actually try asking it questions it is much worse than a bright high school student.
It’s more like someone who trains really hard on many, many math problems, even though most of them are not the replicas of the test questions, and get to that level of performance.
Since the test questions were unseen, the result still suggests the person has some intelligence though.
Note that there’s some transfer learning in LLMs. Training on math and coding yields better reasoning capabilities as well.