Why do you take it personally? Just because you are genuinely interested doesn't mean there are army of people who is sitting their with the sole purpose of gaming leaderboards.
If you think that doesn't happen, I have a bridge to sell you
If you think that doesn't happen, I have a bridge to sell you
Genuine question, because as far as I know there is no feasible way to game this score. But if there is, I want to know.
However that process works would be the thing to circumvent, like by getting it to disclose some letters of its name, or any value that is different between models but the same or similar for each one. I assume it doesn't provide the numerical token values in the output or it would be trivial.
Thank you for clarifying: to confirm, yes, I do understand your claim is a vast cabal of OSS zealots rigs LMSys's blind A/B testing.
I don't think there's anything more of value I can contribute to a discussion on that topic.
Have a good weekend!