Llama 3 8B Llama 3 70B GPT-4
MMLU 68.4 82.0 86.5
GPQA 34.2 39.5 49.1
MATH 30.0 50.4 72.2
HumanEval 62.2 81.7 87.6
DROP 58.4 79.7 85.4
Note that the free version of ChatGPT that most people use is based on GPT-3.5 which is much worse than GPT-4. I haven't found comprehensive eval numbers for the latest GPT-3.5, however I believe Llama 3 70B handily beats it and even the 8B is close. It's very exciting to have models this good that you can run locally and modify!GPT-4 numbers from from https://github.com/openai/simple-evals gpt-4-turbo-2024-04-09 (chatgpt)