This is a limitation of any public benchmark. Anthropic could theoretically identify the prompts and treat them differently, and there’s no way for an external observer to prove that isn’t happening. Which is partly why we need more capable open-source models.
A few things make it harder: the panel contains questions drawn from multiple benchmarks rather than one recognizable test, the evaluation is automated and fixed ahead of time, and the raw outputs/results are public so odd behavior can be inspected.
But ultimately LiveNerf measures the behavior exposed through the API on a fixed public panel. It can’t prove what’s happening internally or guarantee the provider isn’t conditioning on the benchmark.
Longer term, I’d like to add held-out/private or periodically refreshed panels specifically to make benchmark recognition harder. I just don’t want to quietly change the current panel, because having a fixed instrument is important for the longitudinal comparison.