Face recognition, bad people and bad data
ben-evans.com
ben-evans.com
Your not wrong, but I think this is worth taking a second to clarify.
See, when I think of "clean" data, I think of a set of crisp, noiseless images, or text without typos, or DNA sequences w/o sequencing errors or low-quality reads.
For AI purposes, you want a set of representative noise... w/o it, you'll be building a very brittle model. The key is, as you've mentioned, to assure that the method of data collection/generation/processing is either (1) unbiased throughout or (2) is done using various methods ideally performed by different entities.
So, there is certainly a distinction between bias and noise (the opposite of which I would probably call "cleanliness").