Are You Smarter Than An LLM? (Quiz based on the most popular LLM benchmark)
d.erenrich.net
d.erenrich.net
But it can be a good honeypot for either LLMs that cheat or are overfit.
As a result of an accident, Abdul lost sight in his right eye. To judge the distance of vehicles when he is driving, Abdul is able to rely on cues of
A. I only
B. II only
C. III only
D. I and II only
I don't think this is the models learning the specific dataset, but rather the best performing models having learned how to score well on multiple choice tests, such as when not sure to guess an exclusionary combined answer (I'd wager this strategy ends up correct more often than incorrect when all answers seem equally probable based on available knowledge).Which in turn is a useful reminder that we'd best be wary of turning measurements into targets, as we may be targeting adaptations that score well on the targeted measurement but aren't more broadly applicable (and might even undermine a better generally performing model that isn't as smart at acing the test format vs the test content).
What has this to do with being smart? Even Forrest Gump can memorize the answer to this question.