- Apple: “Backup your phone to iCloud, it will be safe there.”
- 5 minutes later: “We’ve wiped your account because of a photos of (porn actor here) which is not CP but technically minor at the time she filmed.”
- “Also we’ve wiped your iPhone because we couldn’t knowingly let you keep that. Good luck contacting your parents, we’ve deleted your contacts. Good luck! PS: We’ve reported you to the police.”
- Also you can’t connect to your iMac now.
We have a Tumblr set up for family to view pics of the kids. Several photos and videos of our kids when they were under 2 were taken down either temporarily or permanently by their CP algo.
These were a pic or video of kids in the bath or without a shirt. In none of them could you see bum or bits. Just a semi naked baby.
Algorithms like this get things wrong all the time
How could this be the case? If it's been determined to be CSAM then it is, by definition, illegal.
If it were true that the database is likely to contain legal material, how would we possibly know about it, given that the contents of the database are secret?
Certain images are CSAM by _context_. They do not necessarily require those within the image to be abused, but rather that the image at one time or another was traded alongside other CSAM.
> If it were true that the database is likely to contain legal material, how would we possibly know about it, given that the contents of the database are secret?
Tools like Spotlight [0] make use of the database, so certain well-known images are known to flag. Such as Nirvana's controversial cover for Nevermind.
[0] https://www.wired.com/story/how-facial-recognition-fighting-...
At the risk of sounding like a broken record, how can we know this is actually true? Every description of the NCMEC database's contents that I've seen is incredibly vague, and as of 2019 it seems like there were fewer than[1] 4 million total hashes available. I would think that if it genuinely did include innocent photos of people's kids, the number would be much higher.
> ...certain well-known images are known to flag. Such as Nirvana's controversial cover for Nevermind.
I've heard this multiples times now, but I've never been able to find any evidence of it actually happening. The only instance I could find was one where Facebook removed[2] that Nirvana cover once for containing nudity.
1. https://inews.co.uk/news/technology/uk-us-collaborate-crack-...
2. https://www.theguardian.com/music/2011/jul/28/facebook-nirva...
Remember, this isn't a porn detector strapped to a child detector.
Step 2: Manipulate pictures so that hash collides with CSAM
Step 3: Get pictures back on targets phone so they get scanned.
I don't have the skills or understanding of how the hashes are created but would this be possible?
• has an iPhone;
• has children;
• took photos of their children which could be mistaken for CSAM by a sloppy reviewer;
• is of sufficiently high importance to justify the effort.
And after that insane effort, all you've done is inconvenience your target for a little while until child safety people investigate your family situation and discover that the photos which got flagged were not actually CSAM.
Immediately after the investigation process discovers the hash fraud, Apple will immediately start delving into exactly how their hash algorithm failed in this instance, improving it to mitigate this exploit. So this target better be worth it!
If this was a plausible exploit, surely it would have already happened to people with Android phones since Google has been doing pretty much the exact same scanning of customer images for over five years. (The only difference with what Apple is now doing is where the hashing is performed—but this makes no functional difference to the viability of your hypothetical exploit.)
[1]: https://www.hackerfactor.com/blog/index.php?/archives/929-On...
[0] https://natmchugh.blogspot.com/2014/11/three-way-md5-collisi...
What you are describing is a second preimage attack-- creating a second input with the same hash as a target.
There is no currently known tractable way to create second preimages for MD5.
Obviously nobody should be using MD5, but it can be useful to understand there are circumstances where it's basically reliable unless you have an extremely sophisticated attacker.
Even though the author says they were 3 million MD5 hashes the second time, the first one he calls them SHA1 and MD5 hashes (even though SHA1 is considered weak too).
I wonder what kind of hashes Apple is planning to use. Will it be whatever is made available to them or will they only accept (what is now considered) secure standards?
Regardless, it's not both. Setting aside how the algorithm was created, it's incorrect to say that an algorithm "created with ML" is itself an ML algorithm.
NeuralHash was so named because it was optimised to run on the Apple Neural Engine for the sake of speed and power efficiency.
The image is not fed directly into the hashing function, like taking an MD5 hash of a file or something.
Rather, the image is first evaluated by a neural net that looks at specific visual details, and has been trained to match even if the image has been cropped or anything like that. The results of the neural net evaluation are what is then input for the hashing function.
This is explained in detail in Apple’s documentation they released with the announcement.