So the theory about why training on training data is unexpectedly inefficient could be because LLMs are "using" the full Causality Chain (using some advanced unknown Physics related to time itself) of our universe/timeline, and so if it tries to train on it's own output that's a "Short Circuit" kind of effect, cutting off the true Causality Chain (past history of the universe).
For people who want to remind me that LLM Training is fully "deterministic" with no room for any "magic", the response to that counter-argument is that you have to consider even the input data to be part of what's "variable" in the Anthropics Selection Principle, so there's nothing inconsistent about determinism in this speculative, and probably un-falsifiable, conjecture.