Reddit has a lot of text on it, sure, but it's definitely not the best model out there, or should we expect language models that think it's appropriate to reply with "r/unexpectedoffice" for every single prompt?
I mean, Wikipedia is certainly better, but Reddit covers some niches of language that Wikipedia will never cover. Wikipedia, Reddit, and Twitter are the three predominant sources of data for large language models, and Twitter has been such a pain for researchers to access for a while already.
What about like, blogs, for instance, as a source of training data? Is the differentiating factor here something like, it's lots of people interacting with other people?
Blogs are hard. You have to manage a large collection of shallow links. People like reddit and Twitter because there are very few deep links. And in the old days you just opened up a pipe and Twitter streamed new tweets directly at you.