How would you game it? I think it is clearly the least gameable leaderboard we have. A more valid criticism might be that you don't like the metric it's measuring, but I think it is a useful metric, though certainly not the only useful metric.
You could ask a question on lmsys, check your server logs for the generated response, go back to lmsys and pick the response that your model generated.
Maybe you could also use a better model for requests from lmsys. E. g. use an unquantized model, disable censorship, etc.
I doubt any of the big players are doing that, but you never know.