“Suppose the government puts a certain drug in the water supply … A couple of conspiracy nuts say it makes your fingers fall off one by one, but the government says that’s ridiculous … However, government employees are all observed drinking bottled water exclusively, and if anyone suggests that government employees might also want to take the completely innocuous drug, they freak out … If by chance you manage to slip a little bit of tap water into a government employee’s drink, and he finds out about it, he runs around shrieking like a banshee and occasionally yelling “AAAAAAH! MY FINGERS! MY PRECIOUS FINGERS!”. At some point you might start to wonder whether the government was being entirely honest with you.”
What is it, exactly, that makes AI-generated text so poisonous to AIs but totally harmless to humans? What is mode/model collapse, and why can it only happen to AIs with too much AI text in their training data and not to, say, human students with too much AI text in their textbooks? The people who know the most about this phenomenon seem much more careful about contamination than they are encouraging us to be.