Examples:
- predict a coinflip: easy to verify, hard to learn
- earn $100: easy to verify, hard to learn
- increase paid subscriptions in an A/B test: easy to verify, hard to learn
I won't get into it, but there are many properties beyond verifiability that are needed to saturate a benchmark.
- earn $100: easy to verify, hard to learn
- increase paid subscriptions in an A/B test: easy to verify, hard to learn
but we both know these examples go against the spirit of my point
also, you are underestimating how short a 10 year time frame is. we are close to self driving, the first neural net image model was in 2013. 13 years is a blink of an eye
The problem is that there is a huge perverse incentive. The intelligence is in the training layer not in the model parameters, but the intelligence is really good at remembering things, so if you let it take the test, it can RL it.