And that's not the sum total of ways that the measure can fail... that's a unique way that comes into being because of the intelligence of the LLMs and other future AIs. All the normal ones are in play too, and perhaps other unique ones as well.
"Gaming" even adds a bit of an adversarialness to the process that isn't necessarily present. Plenty of measures end up "gamed" through perfectly natural attempts to maximize the measure. Someone can be perfectly honestly optimizing for "conversion rate" and not notice that they raised it by lowering the initiation rate more than they lowered the conclusion rate. "But I could account for that by measuring..." would miss the point. There is always a divergence, it only gets more subtle.
This of course also is rather glossing over the difficulty of even defining "ethical" to begin with. Some of what Silicon Valley goes to great efforts to train into their models I consider deeply unethical. Who is right? That isn't going to be answered with "whoever is the most ethical", not even in principle.