I know it's hard to objectively rank LLMs, but those are really ridiculous ways to keep track of performance.
If my reference of performance is (like the vast majority of users) ChatGPT-3.5, I have to first know how Llama 3 compares to that to then understand how that new models compare to what I'm using at the moment.
Now, if I look for the performance of Llama 3 compared to ChatGPT-3.5, I don't find it on the official launch page https://ai.meta.com/blog/meta-llama-3/ where it is compared to Gemma 7B it, Mistral 7B Instruct, Gemini Pro 1.5 and Claude 3 Sonnet.
How does Gemma 7B perform? Well you can only find out how it compares to Llama 2 on the official launch page https://blog.google/technology/developers/gemma-open-models/.
Let's look at the Llama 2 performance on its launch announcement: https://llama.meta.com/llama2/ No GPT-3.5 turbo again.
I get that there are multiple aspects and that there's probably not one overall "performance" metric across all tasks, and I get that you can probably find a comparative between two specific models relatively easily, but there absolutely needs to be a standard by which those performances are communicated. The number of hoops to jump through is ridiculous.