Reddit is my warning flag when AI gives me an answer. If I see Reddit linked, I know to disregard the answer entirely. They may as well be scraping 4chan.
But they do seem to be trying to destroy themselves at the moment.
So a massive chunk of LLMs are Reddit data
What I conflated was that WebText utilizes Reddit links and data (mostly prior to 2023/4) and that I combined WebText and common crawl for the original GPT2 bootstrap into one dataset
I was incorrectly connecting Reddit-mediated WebText pipeline to Common Crawl. So thanks for the correction!