People can argue against this approach for privacy reasons, but I think the false positive argument is a relatively weak one.
There will be many false positives, they will be reviewed by people. When there’s more than a few false positives, you will be investigated by the FBI.
Again, Apple nor any company, have access to the source data, just hashes.
They did add a classifier to iMessage. But it's designed to prevent children seeing any sexually explicit images.[2] There wouldn't be a reason to train it on images of children specifically.
[1] https://www.apple.com/child-safety/pdf/CSAM_Detection_Techni...
[2] https://www.eff.org/deeplinks/2021/08/apples-plan-think-diff...
Perceptual matching is used to sort categories of images. A quick DuckDuckGo will turn up many results. No stretch of the imagination will turn this into a bit for a bit comparison. This is a machine learning algorithm used to categorize images. https://www.ibm.com/blogs/research/2019/10/learning-implicit...
Matthew Green is tweeting about this: https://twitter.com/matthew_d_green/status/14230711866160005... and mentions that it is "preceptial hashing”
9 to 5 Mac article, in which they restate that it is not a classical bit-by-bit hash:
https://9to5mac.com/2021/08/05/scanning-for-child-abuse-imag...
“Perceptual hashing is the use of an algorithm that produces a snippet or fingerprint of various forms of multimedia.[1][2] A perceptual hash is a type of locality-sensitive hash, which is analogous if features of the multimedia are similar.”
https://en.wikipedia.org/wiki/Perceptual_hashing
The goal is to verify a black and white copy of an image is identical to a colored original. Search algorithms want a similar thing so they can validate an image contains a blue car. However, a perceptual hashing algorithm must differentiate between different images containing a blue car while matching a photoshopped copy of the same image.
I would hope most privacy conscious people disable iCloud, but that’s another story.
You can setup secure encrypted backups, but the customer losing the key means losing the back so that’s not what consumer focused companies are going to do. In other words any backup service that doesn’t have big warnings that losing your key loses your backup means they can read your data.
"The mud puddle test: You don’t have to dig through Apple’s ToS to determine how they store their encryption keys. There’s a much simpler approach that I call the ‘mud puddle test’"
[1] https://www.apple.com/child-safety/pdf/CSAM_Detection_Techni...
To be clear each image, the image’s NeuralHash, and a visual derivative are uploaded to iPhoto. This allows for the inspection of the NeuralHash algorithm used which I actually prefer.
The phone isn’t downloading the hash database.
Citation: https://www.apple.com/child-safety/pdf/CSAM_Detection_Techni...
The perceptual hashing is based on AI techniques. “The system computes these hashes by using an embedding network to produce image descriptors and then converting those descriptors to integers using a Hyperplane LSH (Locality Sensitivity Hashing) process.”
The difference is AI classification is based on defining something as say a “Cat” and then the AI spits some association with how cat like the image is. This extracts features from an image then compares lists of features to specific images.
From the PDF:
"The system generates NeuralHash in two steps. First, an image is passed into a convolutional neural network to generate an N-dimensional, floating-point descriptor. Second, the descriptor is passed through a hashing scheme to convert the N floating-point numbers to M bits. Here, M is much smaller than the number of bits needed to represent the N floating-point numbers. NeuralHash achieves this level of compression and preserves sufficient information about the image so that matches and lookups on image sets are still successful, and the compression meets the storage and transmission requirements.
The neural network that generates the descriptor is trained through a self-supervised training scheme. Images are perturbed with transformations that keep them perceptually identical to the original, creating an original/perturbed pair. The neural network is taught to generate descriptors that are close to one another for the original/perturbed pair. Similarly, the network is also taught to generate descriptors that are farther away from one another for an original/distractor pair. A distractor is any image that is not considered identical to the original. "
Image classification on the other hand cares about if the image contains say a stop sign or a trash can. That’s useful for self driving cars etc.
Aka classification you might want to match two different bands playing the same song as identical. Where perception hashing would want them to be classified differently.
It is a perceptual Hash of the images characteristics and perceived continent by the algorithm, that’s the “neural” in neuralmatch.
Apple, for yours has been building-in machine learning dedicated chips into their builds so this shouldn’t affect battery life.
Matthew Green is tweeting about this: https://twitter.com/matthew_d_green/status/14230711866160005... and mentions that it is "preceptial hasing"
9 to 5 Mac article, in which they restate that it is not a classical bit-by-bit hash:
https://9to5mac.com/2021/08/05/scanning-for-child-abuse-imag...
It's fuzzy hash, not ""AI"", based. Cloudflare uses it too, and last time i checked the web is still functional.
https://support.cloudflare.com/hc/en-us/articles/36004610611...
How about we turn the tables and have the complainers suggest a solution. Because every single time an approach to targeted child porn takedowns has been suggested, such as datacenter raids which would not affect as many people, someone is screaming about their privacy.
Child abusers evolve and are very happy if law enforcement doesn't.
Fuzzy means that it takes compression and the like into account, because even if just one pixel out of 20 thousand is different, the hash is different too. Fuzzy hash still recognizes it as the same image, so using an algorithm to alter the color etc. won't work.
That's also true for the no-fly list and the Terrorist Screening Database,[1] yet those are full of false positives. And unlike those lists, CSAM databases cannot be independently verified. To do so would require having the original images, which is illegal.
1. https://en.wikipedia.org/wiki/Terrorist_Screening_Database
So if you're charged on the basis of a fuzzy hash matching, you'd subpoena Apple for the photo in your backup that matched, present it to the court (since it doesn't actually matter if it's CP or not to be admissible), and you win the case.
0. https://www.johntfloyd.com/the-difficulty-with-criminal-evid...
The definition of child porn varies around the world. These systems use the US definition. This is not entirely what you might expect. For example, in the USA the courts have decided that cartoons can be child porn even though no actual children are in the picture. Most of the world does not agree with this, meaning an image can be CP in one place but not another. Is Apple going to enforce the US definitions or the ones where the user actually lives?
In the USA, photos an under-age person takes of themselves can also be considered CP.
What counts as a "child" for sexual purposes also varies around the world. Some countries have a lower age of consent than other places. In some parts of the world the age of consent and the age at which a child stops being a child for CP purposes are different, meaning that a teenager can have sex legally but if they take a photo of themselves doing it, they are trafficking in CP.
Finally, what is actually on these image blacklists? Hardly anyone actually knows because of the third rail nature of CP. Tech firms are often delivered image hashes, not even the images themselves, by third party 'charities' of various kinds and tech workers are - for obvious reasons - not normally given access to the actual pixels. Additionally, appeals from users are invariably ignored because people say "legal issues, it's complicated" and so everyone clams up. If FPs occur there is no way to resolve it and the people who see your appeal, if there even is one, won't be willing to actually look at the image to find out what it was.
It should be obvious how much potential for abuse this hands the people who actually manage these CP databases. Literally any image can be made verboten immediately, without any recourse, and basically nobody will ever find out including the people who shut down the affected users.
No. NeuralMatch was “trained” using 200,000 CP images. “Neural“ is likely a reference to the perceptual matching that it uses. It is not a bit for bit match.
Perceptual matching is a technique used for categorizing images based on characteristics and content.
The algorithm will scan your library containing new information and compare it to what it understands as CP.
The specific subject of child abuse is irrelevant in my commentary. It was a commentary on the general category of AI, used all over, for many things, and more and more every day, but nice try.
Edit: You edited out what I was referring to as I was replying.
Being worried about privacy, establishing a precedent for scanning my data against a government database, and the risk of false positives with such an insanely emotionally charged crime is more than mere complaining.
The onus should not be on me to justify why this shouldn't be done. This is something new and it is perfectly fine to argue against it without needing to provide an alternative.
That being said, my solution is to continue to follow the process that law enforcement is currently using.
If you knew how they approach this you wouldn't be satisfied either
Yes. And now they will evolve by developing a simple system to modify pixels in images when they copy and transmit that will easily defeat this hashing system. The only effect this will have is that moral panickers like you will have got everybody's privacy invaded over your moral panic of the day.
This is, and always is, a game of cat and mouse. Law enforcement is always catching up. They are the cryptologists here. They are never ahead, always behind, because they don't know the new protections peddlers are using until they have been in use and later discovered.
No matter what vector you plug, they will use another, and the game continues (sick game). Maybe divide the image into 32 different quadrants and rearrange them, then put them back in the correct order when viewing through a specific image viewer. I'm sure that would bypass whatever detections they've come up in their fuzzy fingerprinting with as the entire image is now different. By the time they catch someone using this, they'll have already moved on to something different, as they always do.
I will never be ok with warrantless searches of my personal property, no matter the reason or justification or subject, and no matter who it is done by (government or private company). And I say that as a survivor of some pretty horrific shit as a kid to the point I fucking tremble with absolute rage when thinking about it 35+ years later. I would be banned from everything for life if I were to honestly state what I would do with these types of people. The movie "Saw" is tame in comparison. I have no compassion or sympathy for these sickos. But when reading world history, I can absolutely see the importance of "innocent until proven guilty" and Blackstone's Ratio "It is better that ten guilty persons escape than that one innocent suffer." Most of human history was the opposite, and it was brutal and full of literal witch hunts. Are we progressing as a species, or regressing in terms of human rights when it comes to technology?
They don't use simple file hashes to match images, but perceptual hashes. That way they can find modified derivatives of a source image. The problem with this approach, though, is that this is ripe for false positives. Two completely unrelated images can have similar hashes.
If they're using fuzzy matching with perceptual hashes, then the space that false positives can exist in for each perceptual hash is huge.