This disinfo really angers me. That is the exact opposite of what I've read up till now. People talking about "NeuralHash" and being able to detect if the image is cropped/edited/"similar". SO what is the truth?
This disinfo really angers me. That is the exact opposite of what I've read up till now. People talking about "NeuralHash" and being able to detect if the image is cropped/edited/"similar". SO what is the truth?
The whole point of the system is that you get a matching hash after mirroring/rotating/distorting/cropping/compressing/transforming/watermarking the source image. The system would be pretty useless if it couldn't match an image after someone, say, added a watermark. And if the algorithm was public, it would be easy to bypass.
The concern, of course, is that all of this many-to-one hashing might also cause another unrelated image to generate the same fingerprint, and thereby throw an innocent person to an unyielding blankface bureaucracy who believes their black-box system without question.
"Find all images and tag them if they look like this fingerprint" doesn't mean that. It means: "Find all images and tag them if they look 80% like this fingerprint".
Which also means that it will allow governments to upload photographs of people's faces and say: "Tag anyone who looks like this".
Worse, this will allow China to track down more Uyghurs, find people based on guides in the form of images that are spread around to stay safe from the Chinese government, and countries like Saudi Arabia can start looking for phones with a significant amount of atheist-related images, tracking down atheists, and killing them. Because that's what that country does.
https://www.reuters.com/article/us-china-apple-icloud-insigh...
Is this list of hashes already public? If not, seems like adding it to every iPhone and iPad will make it public. I get the "privacy" angle of doing the checks client-side, but it's little like verifying your password client-side. I guess they aren't concerned about the bogeymen knowing with certainty which images will escape detection.
> The main purpose of the hash is to ensure that identical and visually similar images result in the same hash, and images that are different from one another result in different hashes. For example, an image that has been slightly cropped or resized should be considered identical to its original and have the same hash. The system generates NeuralHash in two steps. First, an image is passed into a convolutional neural network to generate an N-dimensional, floating-point descriptor. Second, the descriptor is passed through a hashing scheme to convert the N floating-point numbers to M bits. Here, M is much smaller than the number of bits needed to represent the N floating-point numbers. NeuralHash achieves this level of compression and preserves sufficient information about the image so that matches and lookups on image sets are still successful, and the compression meets the storage and transmission requirements
Just like a human fingerprint is a lower-dimensional representation of all the atoms in your body that's invariant to how old you are or the exact stance you're in when you're fingerprinted... technically Federighi is being accurate about the "exact fingerprint" part. The thing that has me and others concerned isn't necessarily the hash algorithm per se, but rather: how can Apple promise to the world that the data source for "specific known child sexual abuse images" will actually be just that over time?
There are two attacks of note:
(1) a sophisticated actor compromising the hash list handoff from NCMEC to Apple to insert hashes of non-CSAM material, which is something Apple cannot independently verify as it does not have access to the raw images, which at minimum could be a denial-of-service attack causing e.g. journalists' or dissidents' accounts to be frozen temporarily by Apple's systems pending appeal
(2) Apple no longer being able to have a "we don't think we can do this technically due to our encryption" leg to stand on when asked by foreign governments "hey we have a list of hashes, just create a CSAM-like system for us"
That Apple must have considered these possibilities and built this system anyways is a tremendously significant breach of trust.
http://www.fmwconcepts.com/misc_tests/perceptual_hash_test_r...
This DaringFireball[0] article states the goal of the system is to "generate the same fingerprint identifier if the same image is cropped, resized, or even changed from color to grayscale."
So while the fingerprint may be "exact", it's still capable of detecting images which have been altered in some way
[0] https://daringfireball.net/2021/08/apple_child_safety_initia...
But hey, I'm just one of the screeching voices of the minority.
[1] https://www.apple.com/child-safety/pdf/CSAM_Detection_Techni...
"Indeed, Neural-Hash knows nothing at all about CSAM images. It is an algorithm designed to answer whether one image is really the same image as another, even if some image-altering transformations have been applied (like transcoding, resizing, and cropping)."[1]
[1] https://www.apple.com/child-safety/pdf/Security_Threat_Model...
This is the confusion, it's only photos being uploaded to iCloud.
That said, the argument that many people in these threads are making is that they say it's reasonable to scan photos that are uploaded once they're on Apple's servers, they just don't want them scanned while they're still on their phones. In either case, the same photos will be scanned -- ones which are in the process of being uploaded to iCloud -- the disagreement is just about exactly when in said process it's okay to do so. Which seems like a pretty fine distinction to me?
Of course, that's just a nasty way to imply that the images match exactly.