Assuming for a moment that AI-generated content will be ubiquitous enough to impact the sources of a new generation of AI tooling, what is the mathematical limit of this recursion? Is it a Mandelbrot set? A (Douglas) Adams-esque 42? Will it be an expose of ultimate truths? The seed of the singularity? Or a bunch of grotesque amplifications of the worst parts of the human condition? Perhaps all of the above?
Since I don't have the time, I certainly hope some forward-thinking grad student, or suitably motivated genius is experimenting with this now.
In the case of LLM trained on mostly LLM generated data you will see a similar increase in bias of the model and a overfitting to LLM generated data. This might lead to limitations of the performance of LLM in the future.
Academia should still be good to train on. As far as "public-facing" sources, maybe we'll be able to prune away AI generated content somehow?