Llama 3 400B Base / Instruct
MMLU 84.8 86.1
GPQA - 48.0
MATH - 57.8
HumanEval - 84.1
DROP 83.5 - Llama 3 400B Base / Instruct
MMLU 84.8 86.1
GPQA - 48.0
MATH - 57.8
HumanEval - 84.1
DROP 83.5 - Llama 3 GPT 4(Published)
BBH 85.3 83.1
MMLU 86.1 86.4
DROP 83.5 80.9
GSM8K 94.1 92.0
MATH 57.8 52.9
HumEv 84.1 74.4
Although it should be noted that the API numbers were generally better than published numbers for GPT4.That said, I am very curious what OpenAI has in their labs... Are they actually barely ahead? Or do they have something much better that is not yet public? Perhaps they were waiting for Llama 3 to show it? Exciting times ahead either way!
While all the others are catching up and in some cases being slightly better, I wouldn't be surprised to see a rather large leap back into the lead from OpenAI pretty soon and then a scrabble for some time for others to get close again. We will really see who has the momentum soon, when we see OpenAI's next full release.
Llama 3 GPT-4 GPT-4-Turbo* (Apr 2024)
MMLU 86.1 86.4 86.7
DROP 83.5 80.9 86.0
MATH 57.8 52.9 73.4
HumEv 84.1 74.4 88.2
*using API prompt: https://github.com/openai/simple-evalsThere were times when I felt that too, but nowadays I predominantly use turbo. It's probably because turbo is faster and cheaper, but in lmsys turbo has 100 elo higher than original, so by and large people simply find turbo to be....better?
Nevertheless, I do wonder if not just in benchmarks but in how people use LLMs, intelligence is somewhat under utilised, or possibly offset by other qualities.
Since price influences my calculus, I can't say this for sure, but it seems being slightly smarter is not much of an edge, because it's still dumb by human standards. For most non-coding use the smart doesn't make much difference (like summarisation), I find that cheaper options like mistral-large do just as good as Opus.
In the last month I have used Command R+ more and more. Finally had some excuse to write some function calling stuff. I have also been highly impressed by Gemini Pro 1.5 finding technical answers from a dense 650 page pdf manual. I have enjoyed chatting with the WizardLM2 fine-tune for the past few days.
Somehow I haven't quite found a consistent use case for Opus.
I figure we will see some Q4's that can probably fit on 4 4090s with CPU offloading.
I could be wrong too but that’s my understanding. Like float vs half-float.
Although 400B will be pretty much out of reach for any PC to run locally, it will still be exciting to have a GPT-4 level model in the open for research so people can try quantizing, pruning, distilling, and other ways of making it more practical to run. And I'm sure startups will build on it as well.
Still wouldn't be super performant AFA token gen, ~4-6 per second, but certainly runnable.
Of course by the time that lands in 6-12 months we'll probably have a 70-100G model that is similarly performant.