I'm not saying this is easy to do, in fact I'd wager it's incredibly hard: but it seems increasingly obvious that there was no curation, no standards, no approvals process, nothing but a ramshackle mad dash to include as much visual content as possible to train these models on, with zero consideration of whether the people who created it were alright with their content being used that way, if the content was legal to use either by licensing or by being incredibly illegal just by it's own existence. We get more and more stories of the ethical lapses of the massive companies/organizations behind this tech and the resounding chorus from my fellow engineers is seemingly just "well just remove it and let's keep going," or shrugging shoulders "that'll happen, we're trying to do research here."
And fair enough but like, when you're doing research, you don't just run out and grab every chemical you can get your hands on and pour them all in a bucket and see what happens. How is this stuff not being found? What good are these datasets if they include this type of material? How can you trust the results you get from your training anymore once you know what went in?