Today, LLMs learn from a codebase that has mostly insightful comments from well-meaning humans.
In the future, the training sets will contain more and more automatically generated stuff I believe will not be curated well, leading to a spiral of ever declining quality.