SWE-bench with a private final eval, so you can't hack the test set!
In a perfect world this wouldn't be necessary, but in the current research environment where benchmarks are the primary currency and are usually taken at face value, more unbiased evals with known methodology but hidden tests are exactly what we need.
Also one reason why, for instance, I trust small but well-curated benchmarks such as Aider (https://aider.chat/docs/leaderboards/) or Wolfram (https://www.wolfram.com/llm-benchmarking-project/index.php.e...) over large, widely targeted, and increasingly saturated or gamed benchmarks such as LMSYS Arena or HumanEval.
Goodhart's law is thriving and it's our duty to fight it.