this leaderboard seems easily cheated/gamed. once enough eyes are on it it will be worthless
That being said Chatbot Arena is a pretty wide variety of test scenarios. If fine tuning made the model perfect, all the small models would get similar scores to GPT4, which they don't. Essentially it ranks how people believe a ChatBot should respond, rather than just zero shot, 1 shot and COT type benchmarks.
You could ask a question on lmsys, check your server logs for the generated response, go back to lmsys and pick the response that your model generated.
Maybe you could also use a better model for requests from lmsys. E. g. use an unquantized model, disable censorship, etc.
I doubt any of the big players are doing that, but you never know.