Common Corpus: the largest public domain dataset for training LLMs
huggingface.co
huggingface.co
I couldn't figure out their search syntax but both date>2000-01-01 and just 2000-01-01 didn't cough up anything, so I also gravely suspect this data will only be antiquated language usage, too, to say nothing of the objectively horrific OCR job they did