I don't think it is vague in the slightest. Take the most simple examples, how many LLM's have you tested making them? There are stylistic choices pertaining to games that is well beyond a 0/1 reward. Even something as basic as breakout or flappy bird can have wildly different quality between models. Yeah, you could call this animal on a bike benchmarking, but I don't think it is. IMO the problem space occupies an interesting area where you can ignore the pass/fail and focus on the actual level of the model to do something beyond that.
I doubt the OP meant something like creating the whole tech stack for WOW.