And really in reality we do have that actual problem to some degree with the presense of non-blind adults. Actual humans are a mass of mixed clued and clueless with a lot of bad input feeding output feeding other bad input around and around and around, not even counting the legitimate fair differences of opinion.
So it's a problem, but I don't think it's a fundamentally new or worse problem than we already have,and have already always had.
The fix is I don't think there is a fix for that any more than there is for the same thing in humans. There just will always be bad data feeding bad reasoning right alongside the other good data feeding good reasoning. It's probably wrong to ever expect anything else, and fail to operate from that assumption rather than the idea that there might ever be some resolution where we don't have to worry about that.
That is also quite an assumption, it could be that training on the output of better LLMs also reduces this worsening of output. There might even be a tipping point where the LLMs get good enough that training on their output is better than training on the output of humans.
And, as I understand it, one that is already demonstrably false: https://arxiv.org/abs/2306.11644