I think the value is in the examples it provides of language.
Top upvoted comments can filter out the useless information and then it can be trained on actual data and refined.
Except when top voted comments are hivemind approved 'funny' quips/responses, or in reply to exercises in creative writing like half the posts in relationshipadvice, iwantthemanager, nuclear/pettyrevenge, etc
Is this a joke that I'm missing? Top reddit posts are frequently trash filled with misinformation.
Many popular LLMs already include large amount of Reddit comment data which is (usually) cited in their respective papers.