Identifying and eliminating CSAM in generative ML training data and models
purl.stanford.edu
purl.stanford.edu
This article is such an example of this "opportunistic" agenda.
I don’t see this issue going away anytime soon, especially since all that fake csam is still legal and basically unstoppable.
1) LAION only contains links. It's the same links one would find on search engines, or any other large datasets. If they wanted to make this seem less biased, they could have framed as "Identifying and eliminating CSAM in public internet datasets", no need to even mention AI or Machine Learning. But I guess they really want to portray those datasets as some sort of "CSAM database" to tarnish AI reputation in the public eye...
2) You could write this same article about, for instance, search engine indexes or the Internet Archive. Which most likely host such material somewhere there as well. But it doesn't involve AI, so those researchers don't care.
3) The only reason why these links were found and remove is because it's a public dataset. Do they prefer that companies use private datasets, where no one knows what is in it, so no "researcher" can write their little misleading manipulative "think of the children!!" article on it?
Damn, I knew training data was kind of wild-westy, but I didn't know it was that sloppy.
They reached that conclusion then naturally asked, if the model can produce that, how much of it does it need to be trained on to produce it? And they discovered the answer to that is very little.
Which could be for any number of reasons that are far beyond how I understand ML.
What actual difference is there between a machine generating this material or a human being drawing it by hand? It seems like we're living with the "pretty serious consequences" already, and yet, it seems like we've found a way to manage those effectively already.
> if the model can produce that, how much of it does it need to be trained on to produce it? And they discovered the answer to that is very little.
The answer may very well be "none at all," particularly if these systems can create image fragments by inference and a different system or even a human can assemble them.
> are far beyond how I understand ML.
If you have a crappy product, get it regulated, it grants it the imprimatur of credibility and it hamstrings your competitors and startups in the space. An understanding of ML may not be required at all to understand this regulatory situation.
> What actual difference is there between a machine generating this material or a human being drawing it by hand?
Page 8 Section 4 Legal concerns ———
U.S. Code[…]uses a standard prohibiting any visual depiction of CSA that is “virtually indistinguishable” from a minor engaging in sexual conduct. The definition it uses specifically references computer-generated material as being in scope.U.S. Code[…]states that any depiction of a minor that both contains sexually explicit conduct and is obscene can be prosecuted with the same penalties[...]
As such, the current status of CG-CSAM appears to be that prosecution is possible for any representation deemed both graphic and obscene,[…]if that material is indistinguishable. Now that CG-CSAM has reached the point that it may[…] test the application of some of these laws may be tested.
Realistic CG-CSAM also presents obvious difficulties when it comes to prosecutions for CSAM possession[…] in general, the appearance of a child being abused has been sufficient for prosecution[…]alternatives may need to be tested that do not require positive identification of a real-world victim (or worse, that would require such a victim or their family to testify)
> The answer may very well be "none at all,"
You may be correct but they do cite a number of counter-measures that will need to be tested and studied.
The rationale used so far has been that it creates a market for production and revictimized the children. If this is not the case, what are we prosecuting people for exactly?
If the goal is to stop actual human trafficking, then target movement. Possession shouldn't be illegal (stops your pissing-cupid garden statue from being declared CSAM tomorrow, and settles making criminals of shutterbug parents at bathtime), but broadcasting, exhibiting or selling anything to do with it (or tutorials for producing it) should be illegal. I'd even suggest uploading it to cloud/external storage should be illegal.
Distribution allowable only between owner and any direct recipient(s), while both parties are under or over 18 (stops the adult-child grooming issue). Solves the underage-revenge-porn issue as well; Bob15m is on the hook for any forwarding or "losing" of the nude selfie he solicited from Alice14f, lest he face CSAM distribution charges. Either of them possessing it wouldn't be illegal, but it would be a liability if lost or stolen, so both are best served deleting it as soon as possible.
They say every advance in tech starts with pornography. This is how you normalize a culture of responsible data stewardship-- by getting people used to the idea of personal risk associated with keeping data longer than you should.
> This dataset was built by taking a snapshot of the Common Crawl repository, downloading images referenced in the HTML, reading the “alt” attributes of the images and using CLIP interrogation to discard images that did not sufficiently match the captions.
In other words, it was based a random subset of the entire public Internet and there was no human in the loop -- and it's not clear how there could ever be, at the scale this project was operating. (LAION-5B contains 5.8B images.) If anything, it's surprising that the Stanford researchers only found a few thousand CSAM images in the data set.
I somehow got the impression from the 404media podcast that they were the ones who initiated the need for this research, given that they were not allowed legally to dig further into the dataset by themselves, but it's possible they were just given advance access to the report
Anything that ingests a huge amount of data is going to have some undesired content slip through the cracks. The issue with CSAM is that not all of it has been identified. You can query against databases like NCMEC's but that's not going to catch everything. The database also grows all the time -- how do you remove training data from a model you've already built?
The only way this is going to work is on a best effort basis. If the model author does everything reasonably in their power to prevent contamination of their model, there should be amnesty for them in the event something has slipped through.
That would not be such a big problem if our society had not tolerated power-mad moderators who nuked most correct posts and posted correct-sounding nonsense over the last several decades.
protect the kids!!! or something. Can we do a scandal about how many 'extremist' images are in the data-set next too? anti-vaxer, climate denier, nazi, religious extremist propaganda, scientific misinformation. Maybe we're all safer off using corporate models with closed data sets so no one gets any of the wrong ideas.
Resist anyone’s attempt to make everything else you list illegal to express or possess.