Llama 3 70B tied with GPT-4 for first place on LMSYS chatbot arena leaderboard
twitter.com
twitter.com
People seem to be saying that Llama 3's safety tuning is much less severe than before so my speculation is that this is due to reduced refusal of prompts rather than superior knowledge or reasoning, given the eval scores. But still, a real and useful improvement! At this rate, the 400B is practically guaranteed to dominate.
It’s also good to keep in mind that the ELO scores compute the probability of winning, so a 5 point Elo difference essentially still means it’s a 49-51 chance of user preference. Not much better than a coin flip between the first 4 models, especially considering the confidence interval.
That means we truly do have a long way to go. However, it's also possible that user questions don't have enough 'difficulty' to distinguish -- eg, if the proportion of relatively easy user questions is high, then there are fewer opportunities for a better model to distinguish itself.
You'd know that one of the entries was a human because of the latency, but you wouldn't know which one (obviously the chatbot response would be delayed to match).
I guess the problem is people would try to game the system by leaving tells in their response so you'd know that a human wrote it. You could block obvious ones but that would create a meta game of circumventing the block.
Furthermore, you'd have to limit the scope of questioning to topics that the human on the other side is familiar with, otherwise they won't be able to come up with an answer to questions that aren't common knowledge, while an LLM will always happily generate pages of plausible-sounding (even if incorrect) text.
I think a more appropriate benchmark would be subject matter experts interviewing an LLM and another subject matter expert, with clear guidelines regarding the style of writing they're expected to match.
I suspect the future is going to be owned by lots of smaller more specific models, possibly trained by much larger models.
These smaller models have the advantage of faster and cheaper inference.
Before too long we're going to see architectures where a model decomposes a prompt into a DAG of LLM calls based on expertise, fans out sub-prompts then reconstitutes the answer from the embeddings they return.
I wish they didn't include the disclaimer just for the sake of having one, they provided a chat format, so they must have trained _some_ chat, and we can see here its good enough at chat to not act like a pure LLM, or even close.
Labelling it with a disclaimer it doesn't need causes general confusion in scenarios like Phi-2, which had no chat tuning but great scores, so people couldn't understand "why it didn't work".
It's a very much English First model, which is totally different from even GPT 3.5 that works amazing in Brazilian Portuguese.
Who did all those votes? I did like 5 of them...
I'm slightly suspicious someone might be gaming the votes, because billions of dollars of company valuations ride on who has the best AI...
This is the first time there's been a non-private model at the top, and I'd wager top 3.
This isn't OSS either.
Not sure where the jaundiced view comes from. I generally believe the free models tend to be overrated, but I don't think it's a good idea to pretend there's widespread gaming of this. There's a strong self-peasantization streak in LLM discussions, and truth is, no one thing may be "correct", but we can't exclude all data without making the very act of discussion meaningless.
I know the power of Zealotry and Cult in the age of internet and this is just one example of that
That's a lengthy way of saying, contra what you imply, I have practical experience & understand this stuff intimately.
I agree with your opinion re: it's unlikely LLaMA 3 70B beats all private models, but think your way of relaying it, claiming there's widespread gaming of it by advocates for open models, is obviously incorrect, and adds more confusion rather than reducing it.
If you think that doesn't happen, I have a bridge to sell you
Thank you for clarifying: to confirm, yes, I do understand your claim is a vast cabal of OSS zealots rigs LMSys's blind A/B testing.
I don't think there's anything more of value I can contribute to a discussion on that topic.
Have a good weekend!
Genuine question, because as far as I know there is no feasible way to game this score. But if there is, I want to know.
However that process works would be the thing to circumvent, like by getting it to disclose some letters of its name, or any value that is different between models but the same or similar for each one. I assume it doesn't provide the numerical token values in the output or it would be trivial.
https://twitter.com/Teknium1/status/1781328542367883765/phot...
I bet a bunch of people use it for actual life tasks.
More impressively even, Llama 3 8B is approximately tied with GPT-4 (depending on the version), as well as Mistral-Large and Mixtral 8x22B, which is mad for the size.
EDIT: My bad, I was looking at the Overall leaderboard instead of the English one as other commenters. Still quite impressive.
"How can I visualize a decision tree trained using PySpark?"
and it very confidently gave me 4 detailed, step-by-step answers (including code!) that are utterly false and useless. :)
What type of VRAM would you need to run 8B locally?
( Q / 8 ) * B = GB RAM, where Q is the Quant level and B is the model size.
So a Q4 7B model is ( 4 / 8 ) * 7 = 3.5GB RAM (or VRAM).
A non-Q model is 16, so 2 * B.
I believe this is before context, which adds a bit.
https://ollama.com/library/llama3/tags
Llama3 8B fp16 is 16GB, q8 is 8.5GB and q4 is between 4.3 and 4.9GB.
Llama3 70B fp16 is 141GB, q8 75GB and q4 is between 40 and 43GB.
I use the Q5_K_M GGUF, which is >99% the same as the original.
I've seen tests that there is close to no divergence with these quantisations, but it rises steeply going lower:
https://www.reddit.com/r/LocalLLaMA/comments/1816h1x/how_muc...
That said, all other benchmarks so far (including my NYT Connections benchmark) show that both Llama 3 models are exceptionally strong for their sizes.