I'm all for open models, but where do the benchmarks show that?
The official page https://llama.meta.com/llama3/ does not show any comparisons with GPT-4 or Claude Opus
Looking at https://arena.lmsys.org/, Llama-3-70b-Instruct is ranked #5 while current GPT-4 models and Claude Opus are still tied at #1. Meanwhile, Llama-3-8b-Instruct is ranked #14
Would love to be corrected, but either way an article should include sources for these types of claims.