> Indeed, a useful criterion for gauging a large-language model’s quality might be the willingness of a company to use the text that it generates as training material for a new model.
It seems obvious that future GPTs should not be trained on the current GPT's output, just as future DALL-Es should not be trained on current DALL-E outputs, because the recursive feedback loop would just yield nonsense. But, a recursive feedback loop is exactly what superhuman models like AlphaZero use. Further, AlphaZero is even trained on its own output even during the phase where it performs worse than humans.
There are, obviously, a whole bunch of reasons for this. The "rules" for whether text is "right" or not are way fuzzier than the "rules" for whether a move in Go is right or not. But, it's not implausible that some future model will simply have a superhuman learning rate and a superhuman ability to distinguish "right" from "wrong" - this paragraph will look downright prophetic then.