[0]: https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...
[0]: https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...
It's like saying how can evaluating 5 years of performance at work be better at predicting someone's competency than their SAT scores.
https://huggingface.co/papers/2306.05685
This paper makes the argument that...
"Our results reveal that strong LLM judges like GPT-4 can match both controlled and crowdsourced human preferences well, achieving over 80% agreement, the same level of agreement between humans. Hence, LLM-as-a-judge is a scalable and explainable way to approximate human preferences, which are otherwise very expensive to obtain."
So, the Arena could theoretically be automated and achieve similar outcomes. Or at least, it could quickly determine a predicted-ELO for every model, which would be interesting to compare against the human-rated outcomes.
Given the possibility of bias, it would make sense to have the judge “recuse” itself from comparisons involving its own output. Between GPT-4, Claude, and soon Gemini Ultra, there should be several strong LLMs to choose from.
I don’t think it would be a replacement for human rating, but it would be interesting to see.
Phi-2 isn't fine tuned for instruction following yet.
For example, consider my analysis [0] based on observing the progression of Large Language Models (LLMs) in a single text adventure.
[0] https://github.com/s-macke/AdventureAI#evaluation-of-other-m...
-Ask any question to two anonymous models (e.g., ChatGPT, Claude, Llama) and vote for the better one!
-You can continue chatting until you identify a winner.
-Vote won’t be counted if model identity is revealed during conversation.
Edit: I missed the third rule. I wonder how smart their detection is.
Do you really need more than this to know which one you’re going to pick? https://i.imgur.com/En37EJD.png
Avatar doesn’t have humans? Seriously?
Your test isn't checking for instructions, consistency, logic, just one fact which the model you chose may have gotten right by chance. It's fine assuming you only expect the model to fact check and you don't plan to have a conversation, but if you want more than that, it doesn't work very well.
I'm hoping there are votes in there which can reflect those qualities and filtering by conversation length seems like the easiest way to improve the vote quality a bit.
I only make technical (pytorch) questions though.
The Glicko rating system is very similar to Elo, but it also models the variance of a given rating. It can directly tell you a "rating deviation."
https://www.reddit.com/r/LocalLLaMA/comments/17jrj82/new_mic...
> Contains inappropriately sourced conjecture of OpenAI's ChatGPT parameter count from this http URL, a citation which was omitted. The authors do not have direct knowledge or verification of this information, and relied solely on this article, which may lead to public confusion
The URL in question: https://www.forbes.com/sites/forbestechcouncil/2023/02/17/is...
This article was written by Aleks Farseev, the CEO of SoMonitor.ai, who makes the claim with no source or explanation:
> ChatGPT is not just smaller (20 billion vs. 175 billion parameters) and therefore faster than GPT-3