There could be an infinite amount of junk, in fact the amount of junk available online before LLM output was added to the mix was already incredibly huge. At this point it no longer matters because you can take 1 or 2 examples of some knowledge you want to train a new model on and you can then feed it to an LLM with instructions saying to create 10,000 variations on this text while still retaining the factual details, you then filter this list and throw away any output with errors or other garbage and what you're left with are a good number of high quality examples of the information you want to add to your training set. Incredibly, there is new research that shows even this might not be needed at some point because there is evidence that transformers can learn from a single example.
Fortunately there are now AIs that can help with the data cleaning task.
But who monitors the monitor?
At a certain point you just need to take a statistically representative sample of the output for review.
Seems like this is part of the feedback loop that will render AI useless though.
No, it's an effective technique that has been used by Microsoft Research among many others.
you dont need to clean all the junk created. you just need to clean enough to feed the models.