The main issue might be that the training set may get polluted by a lot of AI generated text (content that is not supposed to be there, but is there because AI-generated content is becoming widespread online).
Storage is cheap. With ChatGPT being the preminent system, and with OpenAI recording all of its output, it should be possible to exclude modified answers from the training data.
So interesting that people think facts are the only thing that's valuable. Your comment isn't a fact, does that mean it's worthless?
When looked at individually it's not worthless, but when you are talking about a corpus that significantly lacks facts, then the corpus itself is worthless for any mainstream usage other than a social network text generator.