(Not the OP) What I wonder about is why would performance increase with what you suggest. In my mind, you're proposing getting a language model to generate its own training data. But data on its own is not enough to train a good model. The data has to have some kind of information content that guides the learning algorithm to select a good model. If you feed the training algorithm data that doesn't increase its information about what a good model is, you're not helping it train a better model.
For the record, I tried something like what you suggest when I was doing my Master's. That was back in 2014, and I had to train a classifier for a machine learning class. I was given a training set while a separate validation set and test set were kept private (it was all set up in Kaggle as a private competition). To clarify, the idea was that you trained your classifier of choice on the training set, then labelled the validation set with your trained model and submitted the labelling to get a score that you could use to improve your model. The last day of the competition you had to make a choice and submit a final model, that would be evaluated on the test set, for which you had no information.
The problem was that the training data was not very much. There was more data in the validation set, but the data in the validation set wasn't labelled. So I tried to label the validation set with a model I trained on the training set. And, what would you know. My classifier scored 100% accuracy on the validation set. But when I submitted my trained model on the test set it did much worse, I think close to 60% or so.
Empirically demonstrated then: you can't dogfood a classifier to a better version of itself. When you train a classifier on some data, the classifier learns the underlying distribution of the data, with some amount of error. If you then label new data with the trained classifier and retrain the classifier on its own labelling, you end up multiplying the error.
Btw, that doesn't change with language models, large or small, and it doesn't make a difference whether the model has an unsupervised training step or not. As long as your model is, well, modelling, some unknown true distribution and incurring some error, reusing the trained model to generate new data will generate data with error.
So I don't think what you say can work and I'm curious to understand why you think it will. What are you saying will happen, exactly, if you dogfood an LLM's generations, like you suggest, that will make it improve?