Right I guess I am not familiar how automated Benchmarks for LLM work.
I assumed to decide if an LLM answer was good required Human Evaluation.
Lots of ways to evaluate without humans. Most (nearly all) LLM benchmarks are fully automated, without any humans involved.