In fact, at that scale, they would likely be able to get assistance. EDIT: Stanford themselves acknowledges in the original paper that bulk PhotoDNA access was given.
I'm not saying this is easy to do, in fact I'd wager it's incredibly hard: but it seems increasingly obvious that there was no curation, no standards, no approvals process, nothing but a ramshackle mad dash to include as much visual content as possible to train these models on, with zero consideration of whether the people who created it were alright with their content being used that way, if the content was legal to use either by licensing or by being incredibly illegal just by it's own existence. We get more and more stories of the ethical lapses of the massive companies/organizations behind this tech and the resounding chorus from my fellow engineers is seemingly just "well just remove it and let's keep going," or shrugging shoulders "that'll happen, we're trying to do research here."
And fair enough but like, when you're doing research, you don't just run out and grab every chemical you can get your hands on and pour them all in a bucket and see what happens. How is this stuff not being found? What good are these datasets if they include this type of material? How can you trust the results you get from your training anymore once you know what went in?
Just as we can't police a society to the extent it's completely free of crime, otherwise we will not function as a society and/or we will find ourselves living in a totalitarian state?
Also, the totalitarian analogy is total nonsense.
Is that correct?
If so, why are you not calling for the internet to be shut down? The collection under discussion is just a subset of the internet.
You'll also want to either ban or enforce review of all camera output, including closed-circuit cameras. Someone is going to have to watch all those watchers pretty closely...
As for totalitarianism, well, how are you going to make any of this happen? As we just saw you'll need client-side visibility into everyone's visual data stores. Not to mention old photo albums world-wide contain images that would set off ImageNet. There are famous paintings and statues that would, too.
[1] Do you exempt NCMEC? If so, how do you justify that?
Get real. Scraped data is going to have problems. The expectation should be that reasonable measures are taken to filter proactively and that reactive measures are taken to remove illicit content found after that.
> How can you trust the results you get from your training anymore once you know what went in?
Before taking a moral high ground, please take your time to learn how it works. LAION is well known for its notoriously poor quality simply due to being humongous. There's all sorts of unwanted crap in here. Everyone training on it knows that and does filter it, because it's practically unusable otherwise.
blank stare
There are something like ~42000 traffic fatalities in the United States every year, we do not adopt draconian measures in dealing with it, even though loss of life and severe life-altering injuries occur. And such traffic accidents are absolutely devastating for the victims and their families.
Unless the dataset is particularly severely contaminated, the known illegal images should be removed and the rest of it should be allowed to stay up.
Have you considered that maybe you shouldn't?