Either the eval maintainers need to be given the closed source models (which will likely never happen) or the model authors need to be given the private evals to run themselves.
The sport was tainted before Lance and still is.
https://www.anandtech.com/show/15703/mobile-benchmark-cheati...
Given that the models are released to public, the test maintainers can just run the private tests after release, either via the prompts or via an api. Cheating won't be easy.
The models of Open AI, Claude and other major companies - are all available either for free or a small amount(200$ for OpenAI Pro). Anyone who can pay this, can run private tests and compare scores. So, the public does not need to rely on benchmark claims of OpenAI based on its pre-release arrangements with test companies.
Yes, by uploading the tests to a server controlled by OpenAI/Anthropic/etc
The other comment speaks to training on private questions, but training on public questions in the right shape is incredibly helpful.
Once upon a time models couldn't produce scorable answers without finetuning on the correct shape of the questions, but those days are over.
We should have completely private benchmarks that use common sense answer formats that any near-SOTA model can produce.