Laion Took Down Its Datasets After Discovery of Child Sexual Abuse Material
404media.co
404media.co
I'm not saying this is easy to do, in fact I'd wager it's incredibly hard: but it seems increasingly obvious that there was no curation, no standards, no approvals process, nothing but a ramshackle mad dash to include as much visual content as possible to train these models on, with zero consideration of whether the people who created it were alright with their content being used that way, if the content was legal to use either by licensing or by being incredibly illegal just by it's own existence. We get more and more stories of the ethical lapses of the massive companies/organizations behind this tech and the resounding chorus from my fellow engineers is seemingly just "well just remove it and let's keep going," or shrugging shoulders "that'll happen, we're trying to do research here."
And fair enough but like, when you're doing research, you don't just run out and grab every chemical you can get your hands on and pour them all in a bucket and see what happens. How is this stuff not being found? What good are these datasets if they include this type of material? How can you trust the results you get from your training anymore once you know what went in?
Just as we can't police a society to the extent it's completely free of crime, otherwise we will not function as a society and/or we will find ourselves living in a totalitarian state?
Also, the totalitarian analogy is total nonsense.
Is that correct?
If so, why are you not calling for the internet to be shut down? The collection under discussion is just a subset of the internet.
You'll also want to either ban or enforce review of all camera output, including closed-circuit cameras. Someone is going to have to watch all those watchers pretty closely...
As for totalitarianism, well, how are you going to make any of this happen? As we just saw you'll need client-side visibility into everyone's visual data stores. Not to mention old photo albums world-wide contain images that would set off ImageNet. There are famous paintings and statues that would, too.
[1] Do you exempt NCMEC? If so, how do you justify that?
Get real. Scraped data is going to have problems. The expectation should be that reasonable measures are taken to filter proactively and that reactive measures are taken to remove illicit content found after that.
> How can you trust the results you get from your training anymore once you know what went in?
Before taking a moral high ground, please take your time to learn how it works. LAION is well known for its notoriously poor quality simply due to being humongous. There's all sorts of unwanted crap in here. Everyone training on it knows that and does filter it, because it's practically unusable otherwise.
blank stare
In fact, at that scale, they would likely be able to get assistance. EDIT: Stanford themselves acknowledges in the original paper that bulk PhotoDNA access was given.
There are something like ~42000 traffic fatalities in the United States every year, we do not adopt draconian measures in dealing with it, even though loss of life and severe life-altering injuries occur. And such traffic accidents are absolutely devastating for the victims and their families.
Unless the dataset is particularly severely contaminated, the known illegal images should be removed and the rest of it should be allowed to stay up.
Have you considered that maybe you shouldn't?
How's this fundamentally different than Google (or any other search engine) indexing a site that contains CSAM? It's not technically feasible to filter every image on every indexed page during crawl time, especially if it requires uploading every image to PhotoDNA. The service would simply break under the load.
There's likely images of animal and human torture and mutilation in there. And there are people that fantasize about doing such things all the time. Nothing done about that, though.
Like trying to ensure all my kitchen surfaces are 100% free of germs. Not possible, or the amount of effort and disruption caused by the cleaning will make it impossible to get on with my life. Nasty stuff exists out there. We don't want too much of it, though. That's my opinion about it.
With open source efforts, we can eventually reach a state of public certification. It can be fully vetted, unlike closed source models.
like right now, if someone tries to download old laion and find those images manually, it would probably take them a very very long time. if they had the diff, they would find them in a few minutes.
i dont think its a good argument, but its an argument. the researchers found them without a diff so likely someone else can too. so theres not much reason to leave them in that i can tell.
Like, "identified by MD5" doesn't mean that they didn't attempt to remove images! A lot of CSAM out there is not yet on those lists, and a lot of it is being added as we speak, so if LAION-5B was filtered using one list of MD5 hashes, that doesn't mean it doesn't contain CSAM that's on a different list (or a newer version thereof).
And you have this problem no matter whether LAION actually took any steps to remove CSAM before publishing the dataset or not. Unfortunately, any sufficiently large such dataset will contain some bad content (how we should deal with this is a different question that I sadly don't have an answer for).
https://www.exploit-db.com/docs/english/46047-md5-collision-...
Get the dataset minus the illegal data back. It's essential to maintain parity with closed source efforts.
I'd bet that OpenAI also has this in their dataset too. And nobody will ever even know. Open source is the only AI that can be knowingly scrubbed and cleaned.
there is no panic here. every reasonable person agrees this is a problem, and it will take very little effort to fix it. these AI companies are not hurting for money. every single one of them is valued in the billions of dollars. they can spare "0.5 ppm" of their funding to go do what almost everyone agrees needs to be done.
I'm not saying that more shouldn't be done, but that doesn't mean that the solution is easy or that a perfect solution even exists.
This makes me wonder if CommonCrawl WARC dumps themselves contain CSAM.
So this is effectively preventing the creation of any new models.
E.g. one of the issues with changes in Twitter moderation after the management change was that it turned out that reducing the moderation manpower meant that suddenly CSAM on Twitter was more prevalent.
The same applies to other media - e.g. observing Mastodon https://www.theverge.com/2023/7/24/23806093/mastodon-csam-st... found a pretty similar "CSAM rate" as in this Laion dataset.
What about cases where one man’s cute family photos in the bathtub are another man’s pornography? There’s a bunch of reasons why that’s impossible.
Are there statistics on child pornography over time? Is it spiralling out of control? Is it diminishing to zero? It would be easier to live with the pearl clutchers if there was lots of good evidence to show that owning illegal JPEGs causes either JPEG production or abuse offender rates to increase, or that decreasing JPEGs correlated with a downturn in offending.
It also might be ok to ban stuff because it feels like the right thing to do, but it’s not an amazingly solid argument.
There are thousands of images of real murder[1] and other absolutely heinous acts committed against humans and animals out there. Millions and millions of murder fiction books purchased every year. So much violence in the media. And people with secret obsessive interests and fantasies of torture and homicide, who view large numbers of such images intentionally. And a large number of homicides each year.
We do not prosecute people for looking at images of murder on the Internet.
So comparatively little is done to stop that, instead we focus on sex-related moral panics, driven by what is likely a primitive primal human instinct.
1. War crimes, etc.
You’re not legitimately comparing the puritanical sex-averse aspects of the US and child porn, are you?
I think nobody should be going to prison for looking at or storing any image on their computer. Or reading any book or text as well (some countries prosecute possession of certain written material). Such actions are part of the private sphere, which the state has no business policing, unless you commit some violent crime or are severely insane, for example.
Sending people to prison for that absolutely is a moral panic in my opinion. It is a modern day form of heresy or witch hunt.
The laws regarding child pornography possession are "strict liability" in many jurisdictions. So these laws prevent entirely legitimate activities from being carried out, because of the fear of being persecuted (it has parallels to the Salem Witch Trials).
Can you name a single legitimate activity that is prevented from happening due to the need to comply with CSAM related laws?
and as a society we dont need statistical evidence for every single law we come up with. some things we just know are harmful to victims because of common sense.
Heck even we still have problems with society discriminating against young people, and especially teenagers, by the way. So I really care about young people's rights.
But I still think these draconian content laws are unacceptable, they are replacing one form of abuse (pedophilia), with another by the state (sending people to prison for possession of data).
if we are going to argue over historical injustice, this crime in particular has been ignored for centuries because the perpetrators were often powerful and wealthy.
Society and AI development have to keep going, otherwise progress will be disrupted. It's progress that has gotten us out of the barbaric Medieval period and the Dark Ages.
Doesn’t this mean that the crime is in the act of publishing? If they have the photo but don’t publish it then what have they done wrong? What about if the photo is a very realistic rendering of a child that doesn’t exist? (Or just a drawing, for which people have indeed been prosecuted in some parts of the world.)
(The flip side is already covered in law: if the face is real but the scene is fake then it is still an invasion of privacy.)