This repo already has too much visibility now. Anthropic will soon benchmaxx it.
A few things make it harder: the panel contains questions drawn from multiple benchmarks rather than one recognizable test, the evaluation is automated and fixed ahead of time, and the raw outputs/results are public so odd behavior can be inspected.
But ultimately LiveNerf measures the behavior exposed through the API on a fixed public panel. It can’t prove what’s happening internally or guarantee the provider isn’t conditioning on the benchmark.
Longer term, I’d like to add held-out/private or periodically refreshed panels specifically to make benchmark recognition harder. I just don’t want to quietly change the current panel, because having a fixed instrument is important for the longitudinal comparison.