suppose you have a partner who is a 'petite' woman of 34. She enjoys posting nudies on a website, but without her face in that picture. Someone who collects child porn downloads it, because he enjoys that picture. A year later he gets caught by the police and all his pictures get marked as 'verified child porn'. Suddenly you get marked as owning child porn.
https://www.nytimes.com/interactive/2019/09/28/us/child-sex-...
Apple’s thing has some sort of threshold anyway so one image would not trigger it. I don’t buy your example - the CSAM images are not what you’re describing.
They are instead fuzzy classifiers, and thus have non-zero error rates.
How can you do that, considering md5 can have collisions?
For a hash (whether cryptographic or perceptual), there is a chance of random collisions and also a difficulty factor for adversarially-created intentional collisions. The random collision probability has to be estimated based on some model of the input and output space (with cryptographic hash functions, you would usually model them as pseudorandom functions and assume that the collision probability is the same one created by the birthday paradox calculation).
Intentional collisions depend on insight about the structure of the hash function, and there are also different kinds of difficulty levels depending on the nature of the attack (preimage resistance, second-preimage resistance, and collision resistance). Gaining more insight about the structure of the hash function can act to reduce the work factor required for mounting these attacks. That should be true for perceptual hashes just as much as cryptographic hashes, but presumably all of the intentional attacks should start off easier because the perceptual hashes' threat models are weaker and there's much less mathematical research on how to achieve them.
And in AI systems involving classifiers, it was generally easy for people to create adversarial examples given access to the model. Perceptual hashes for estimating similarity to specific known images aren't the exact same thing because it's less like "how much like a cat is this image?" and more like "how much like NCMEC corpus image 77 is this image?", but maybe some of the same techniques would still work. In the cryptographic hash analogy, I guess that would be like trying to break preimage resistance.
To mitigate adversarial false positives one idea is to use the combination of a cryptographically strong hash along with a randomly selected perturbation of the file. Prior to hashing, perturb the file and submit both the hash and the selected perturbation to apple. Apple selects the DB based on the perturbation and proceeds with matching and thresholding.
If the attacker does not know how the image will be perturbed prior to hashing then he cannot generate an image which matches with known CSAM.
I think the rest of us have been discussing how this can and will be abused, by definition by adversaries.
Many of us have also observed for years how systems are abused so we sadly have a gut feeling for this.