Maybe this can help
https://simonwillison.net/2024/Oct/25/pelicans-on-a-bicycle/
https://simonwillison.net/2024/Oct/25/pelicans-on-a-bicycle/
The best LLM benchmarks test around the margins of those behaviors, tasks that are difficult and correlate with usefulness while being removed enough to stay unpolluted