How do you see a possibility of avoiding model collapse? It seems to me to be an unavoidable consequence of two facts I consider pretty much irrefutable:
1) LLM content can at best be as good as the source material it was trained on. That's the upper bound. "Out of distribution" output of LLMs is mostly unusable.
2) LLM-generated content increasingly drowns out original content, online and elsewhere.
This spells "monotonically decreasing content quality" to me.