Exploitive, illegal photos of children found in the data that trains some AI
washingtonpost.com
washingtonpost.com
I can't find the actual study, but I am wondering if anyone actually verified the photos are CSAM, and not a collision or even just misidentified images?
There is substantial evidence that the technique used is highly likely to produce false positives when ran on a large dataset like this.
The study says 825 were confirmed using one method, and 183 with another- not clear to me if those are unique or duplicates of each other. Also, "confirmed" only means "is or is likely" CSAM, and they don't provide any numbers as to how many were classified as "likely". It also isn't clear how they define "likely".
EDIT: Previously I said the article was overstating the number of CSAM images found, but I since re-read it and the article may be citing the correct total from the two methods)
Why would a collision happen here?
Any minor alteration to an image will change it's cryptographic hash completely, making those pretty poor tools for finding images that might want to be hidden. Perceptual hashing should match e.g. resized versions of the image, but with them two "blurry thumbnail hashes" can match accidentally.
https://www.wsj.com/articles/chatgpt-openai-content-abusive-...
I know some jurisdictions have laws to the point that a drawing of two stick characters having sex with the words "Lisa" and "Bart" on them would be illegal, so presumably an AI generated image like above would also be illegal, but I don't believe that's a uniform law across the world.
The accessibility of CSAM content links within the dataset means those images are already on the internet being made publicly accessible by another website, which was found during the LAION image crawl.
(interestingly this does rather imply that if LAION was able to compare images to the relevant CSAM hashes during the crawl then they'd simply be able to exclude them automatically - although I imagine that database is always growing)