Another thing that sort of puzzles me about benchmarks is that LLMs are not deterministic and do not always complete a problem. So what are the results actually representing? The best run? The average? It is all in some ways a falsehood
Generally though, many of these fairness complaints do go away if there is "3rd party testing". Right now, companies reporting their own benchmarks has all the problems that 3rd party testing resolves in many other industries.