I know the power of Zealotry and Cult in the age of internet and this is just one example of that
I know the power of Zealotry and Cult in the age of internet and this is just one example of that
That's a lengthy way of saying, contra what you imply, I have practical experience & understand this stuff intimately.
I agree with your opinion re: it's unlikely LLaMA 3 70B beats all private models, but think your way of relaying it, claiming there's widespread gaming of it by advocates for open models, is obviously incorrect, and adds more confusion rather than reducing it.
If you think that doesn't happen, I have a bridge to sell you
Genuine question, because as far as I know there is no feasible way to game this score. But if there is, I want to know.
However that process works would be the thing to circumvent, like by getting it to disclose some letters of its name, or any value that is different between models but the same or similar for each one. I assume it doesn't provide the numerical token values in the output or it would be trivial.
Thank you for clarifying: to confirm, yes, I do understand your claim is a vast cabal of OSS zealots rigs LMSys's blind A/B testing.
I don't think there's anything more of value I can contribute to a discussion on that topic.
Have a good weekend!
https://twitter.com/Teknium1/status/1781328542367883765/phot...