GPT-3 was 175B, models like Gemma4 with 31B vastly outperform it, so there is more to it
as Karpathy noted, the initial GPTs were trained on complete garbage (literally, the average document from the Common Crawl is random nonsense), yet they worked. now we can use present LLMs to curate the data for the next generation