That's a good point and conventionally if benchmarks aren't run as "one shot", it is denoted as "benchmark@K". Inference time scaling has historically shown improvement.
Generally though, many of these fairness complaints do go away if there is "3rd party testing". Right now, companies reporting their own benchmarks has all the problems that 3rd party testing resolves in many other industries.