Do newer coding models end up training on the AI slop generated by older models?
As old coding models were writing and pushing code to public, does the new coding models train on them because generated code by old coding models were not so good?
In reality there are lots of ways to bias towards quality and all decent labs will be feeding quality signals via techniques like RLHF and supervised training so that you don't get the median human-slop output, but the borderline super-human modern LLM output.
I'd argue that the way attention works plays into the slop too