Yeah but thats literally above ASI, let alone AGI.
Average human scores <1% on this bench, opus scores 97.1% when given an actual vision access, which means agi was long ago achieved
New benchmark idea:
20 questions of guess the number 1-10, with different answers.
We run this on 10,000 humans, take best score.
Then we take 50 ai attempts, but take the worst attempt so "worst case scenarior robustness or so".
We also discard questions where human failed but ai passed because uhhh reasons...
Then we also take the final relative score to the power of 100 so that the benchmark punishes bad answers or sum.
Good benchmark?