Common Crawl is heavily populated with reddit threads
So a massive chunk of LLMs are Reddit data
So a massive chunk of LLMs are Reddit data
What I conflated was that WebText utilizes Reddit links and data (mostly prior to 2023/4) and that I combined WebText and common crawl for the original GPT2 bootstrap into one dataset
I was incorrectly connecting Reddit-mediated WebText pipeline to Common Crawl. So thanks for the correction!