The SEO spam industry will just pivot to infect the results of those language models.
That's something many humans already suffer from, due to social media platforms creating large echo chambers.
I think there's enough human-generated content that LLMs won't suffer from that. If they did, you could always manually filter what they train from... only data from websites with trustworthy timestamps pre-2020, for instance, or content after that which might be partially AI generated but still has strong human filtering, like wikipedia and arxiv and scientific papers in general.