The rate at which AI can generate text will be so much greater than what humans can generate.
You're already assuming pagerank and upvote systems won't break down in the future.
Presumably at some point computers will become (already are for all I know?) the largest consumers of content on the internet as well as its producers.
Now go ahead and spend $50 dollars on AI generated text nobody is ever going to read, just like almost nobody is going to read this comment.
Note that GPT-3.5 and above are already intentionally polluted with their own output by the RLHF process.
i'd say llm's represent a institutionalized reinforcement of bias (much like journalism) combined with some in-human autonomy.
So, not like a watermark, which would be impossible.
But: very fragile, especially if people are specifically trying to hide their GPT use, or have access to the watermarking algorithm or online oracle.
And: other methods – like remembering all output ever, or fuzzy summary representations of all output ever – seem to me similarly fragile, & introduce other problems & impracticalities.
A guess: OpenAI internally initially shared the common concern that "consuming its own junk outputs" could be a problem. But their own experiments so far, private & public, may have convinced them it's not as much of a problem in practice as it seems in theory. The model outputs have a mix of good and bad text – just like the pre-LLM internet. And, the same filterings/weightings that have worked on pre-LLM content keep working. And, counter to some early intuitions, often one LLM's quality output is in fact very-useful input for other later LLMs.
(You can skip to the section “Verifying“)