This was pretty much refuted by Meta with their LLama3 release. Two key points I got from a podcast with the lead data person, right after release:
a) Internet data is generally shit anyway. Previous generations of models are used to sift through, classify and clean up the data
b) post-processing (aka finetuning) uses mostly synthetic datasets. Reward models based on human annotators from previous runs were already outperforming said human annotators, so they just went with it.
This also invalidates a lot of the early "model collapse" findings when feeding the model's output to itself. It seems that many of the initial papers were either wrong, used toy models, or otherwise didn't use the proper techniques to avoid model collapse (or, perhaps they wanted to reach it...)
So yes we will see manual labor to finetune the data lair but this will only be necessary for a certain amount of time. And in parallel we also help by just using it: With the feedback we give these systems.
A feedback loop mechanism is fundamental part of AI ecosystems.