Edit: Downvotes because?...
For example when CIA/NSA tools leaked, one of them had precisely this purpose.
Conversely, the governments of the world get to keep trying, over and over.
It's significantly harder to develop collisions for an algorithm you cannot inspect or obtain the output of.
(well also, emailing actual CSAM is way easier and mostly just gets the sender reported)
You're exposing the hashes to the world, you're not able to E2E encrypt anything if you need to server side scan, which you probably do since trusting the client no matter what is generally bad in potentially adverserial situations, and you get all this negative press and loss of reputation and potentially pressure from governments around the world to use this for other ends. Cui Bono?
The second scan applies only for those images which are flagged as positive, which are then accessible by Apple. This is applied to detect adversarial hashes. The rest of the images stays encrypted. So, indeed on-device scanning is the only way to enable at least partial E2EE with CSAM detection.
Yes, this was PR failure Apple. They rushed the announcement because of the leaks, and secondly they thought that people will understant the system when they did not. There is too much misundersting. That scanning for example is built-in so deep into the iCloud pipeline, that one does not simply change the policy for scanning the whole phone.
On a technical level I think you're correct. As a holistic approach to the problem, I still disagree. This is too cute for its own good. The PR misunderstanding is a symptom of that.
>The second scan applies only for those images which are flagged as positive, which are then accessible by Apple.
In the end, Apples software is scanning all of the images, why is it any more privacy respecting to do it this way? I guess reasonable people can disagree on that, personally I wasn't fully aware of the cloud side scanning either, and I don't think the public was either. This is similar to Snowden's revelations, if you were paying attention you probably already knew a lot of that, but the incident made everyone aware of it in a very blunt way.
>The rest of the images stays encrypted
I think this is unclear, Apple can still decrypt those other images, how else could you view them in a browser?
This goes back to what Stratechery said about capability vs policy.
Obviously there is change coming to iCloud. Otherwise whole PSI protocol is pointless.
And, to at least respond to two obvious counter-arguments I've looked into:
"But it's just one line of code to change it to scan images (that aren't uploaded to iCloud) (that are anywhere on the device)!" No, it isn't; if you read the technical documentation and the more technically-oriented interviews Apple's given on this, there isn't just one hash that needs to be matches, there are two hashes, one on the device and one on the server. (I think Apple did a very poor job of communicating this to a general audience; it certainly wouldn't have alleviated all the concerns, but if it was understand as "client-server scanning" rather than "client-only scanning" it might have at least changed the tenor of the conversation.) That doesn't mean they can't do a combination client-server scan on every single image or even file on the device, but it makes it both more difficult to do and more difficult to hide.
"But what if the system doesn't work as Apple's described it?" Well, if you don't trust Apple to some degree, all bets are off. They already do ML-based image analysis of all photos in your photo library regardless of iCloud status and they've literally demoed this on stage during iPhone keynotes, so if Apple was going to secretly give government access to on-device scanning, a different technology -- one that works (questionably well) on all images, not just already-learned ones -- is literally already there. The only way you "know" what Apple is doing with your data on or off device comes from a combination of what they tell you and what third-party security researchers discover.
If they get caught doing secretly this in China that would be a big blow. But if they are doing openly and it is known that the government provides the hash, they can wash their hands.
Tangentially, I think keeping their servers clear of illegal material is actually Apple's main motivation. This, in turn, supports claims made by nay-sayers that Apple could scan for other types of images/content in more repressive countries (but not necessarily report the people who did it). However, this assumption also contradicts arguments that Apple will start scanning for pictures of (e.g.) drugs, or weapons. Such images are not inherently illegal and therefore of no interest to Apple.
My guess about this whole debacle is that - with pressure from the government to scan their cloud storage - that this is the alternate scenario to avoid giving up (or being forced to) the "encryption" guarantees of their cloud. I'm not sure what technical process they have in place to "only decrypt with valid law enforcement requests" or allow account rescue, but it seems likely that not just any employee can view whatever they want, before or after this system.
Saying that I can maybe see a way the pressures are on this doesn't mean that I'm saying this is a good solution though. Clearly technically implementing this is opening a can of worms that can't really be closed again and makes a lot of other scenarios "closer".
Also, evidently, people are a lot more comfortable with the idea of them actively scanning stuff people store in the cloud than transmitting the information in a side channel so they don't even need to handle decrypted data without a hit.
This is exactly the case. https://www.eff.org/deeplinks/2019/12/senate-judiciary-commi...
> I'm not sure what technical process they have in place to "only decrypt with valid law enforcement requests" or allow account rescue, but it seems likely that not just any employee can view whatever they want, before or after this system.
They have master keys that can be used to decrypt almost everything you upload. They can be compelled to decrypt and turn over information on anyone. Another (unsourced) comment in this thread indicated they do so 30,000 times per year. The new encryption scheme will effectively stop this for photos, and no doubt other files in the future.
Apple's side will begin using shared key encryption, which will require ALL ~31 keys to decrypt the offending images.
The decryption keys are only generated on-device, and only from a hash that results from a CSAM match. The other photos, simply won't have the decryption keys generated, so they don't even exist.
As an interesting side note, this means that a person who surpasses the CSAM threshold will still only reveal the images that actually match the CSAM database. Every other image, including those that have CSAM unknown to authorities remain encrypted. This is hardly a big win for the big scary government. They now have far less ability to search for evidence of any other crimes. You could upload video of your bank robbery to iCloud, and as long as your personal device remains secure, nobody will know.
Even if that was a goal (and I would argue they have a hard stance against it), this system as built is not usable for that.
While they can scan locally, every step of recording, thresholds, and subsequent automated/manual auditing is built to require content to be uploaded to iCloud Photos.
Even assuming that is true that iCloud is trivially “hackable” - and as I understand it, that was never clear how those leaks happened - how does uploading to iCloud help when it specifically needs to be uploaded from the users phone along with the scanning metadata.
In fact, isn’t apples proposed implementation here the _only_ cloud service that protects against your proposed attack - while other clouds scan stored data and can be triggered by your attack, Apple’s requires you to upload specifically from a registered phone on their account; data stored on-cloud is never scanned.
The response to this is “yeah but then a human will review it and nothing will happen to the victim of the attack, because it’s just some slightly blurry ordinary images”. But it ignores to entirely likely harms that could result from that.
1) Law enforcement use it as the basis of getting a search warrant, but conveniently leave the bit about the alerts being false alarms off the warrant application.
2) The list of people who have had CSAM alerts is inevitably leaked to the public, and the victim has to spend the rest of their lives explaining to people like employers why Apple flagged them as possessing child sexual abuse material.
At the end of the day, all the gaslighting about “no potential for inadvertent harm” is bs, because it’s my device, so get lost. Go run your anti-privacy software somewhere else, imo.
- Law enforcement doesn't get _anything_ unless it triggers a large number of images that match the perceptual hash
- It also needs to match a -private- perceptual hash, that isn't distributed to devices, and so we don't have a reliable way of generating collisions for or even knowing that collisions are generated
I mean this whole thing is bad enough on its own without having to artificially manufacture extremely specific scenarios and extrapolating from there to invent hysteric conclusions.
But you’re right though, the possibility of the list being leaked and ruining countless innocent lives is the much more likely of the two scenarios I described.
In that case, nothing is stopping them from scanning everything already uploaded anyway, and nothing is stopping them pushing code to your device to scan it without telling you about it. Nothing is stopping them or the government from making these "lists" anyway.
I'm not saying you (or anyone) should trust Apple, but if you already don't - then this changes literally nothing.
Feel free to respond to my point about the alert catalog inevitably being leaked and ruining lives. Or you could just have a go at gaslighting me a little more if you prefer.
- Send colliding images
- image gets uploaded to icloud automatically
- image _also_ collides with private hash <- completely unclear how this happens
- Only the colliding images are looked at by apple and are determined to be innocent
- user goes on a list (This is an imagined scenario)
- User is reported to law enforcement even though the images are innocent (This is an imagined scenario)
- Law enforcement uses this hypothetical report to file a warrant (This is an imagined scenario)
- Law enforcement uses the hypothetical warrant to extract images that are completely innocent, and somehow build a case around this
- The "List", which is entirely a hypothetical of yours, "leaks" (This is an imagined scenario)
Which also requires:
- Apple does not counter the meaning of "the list"
- Apple is not sued for vast quantities of money
I expect the first argument is that none of this matters as long as "the idea" is out there, the reputational damage is already done. Except if that's true, then none of this is necessary at all, just make the accusation.
So, sure, continue to thread the needle between "They are automatically sending all information to the government, so promises are meaningless" and "This new process, on top of them potentially sending all information to the government, somehow makes it worse".
I mean, this is all a million times more difficult and less likely than just, like, sending them CP in the first place. Or uploading it to their Gmail or any other cloud they use. Or just send a report that they have it to the police without actually doing anything.
All it requires is somebody to send somebody else a colliding image.
This will send an event to Apple. There is nothing imaginary about that, it is exactly how the system works.
Now that Apple has this information, the only thing left is for it to be leaked or compromised in some way.
This is much simpler than the scenario you’ve described, because they require the attacker to first commit the crime of possessing CP. Its also possible to do without tipping off the victim in any way.
Apple, in case you didn’t know, is a company that had already been the source of a couple of the most notorious data breaches ever (and has somehow managed to so far avoid getting “sued for vast quantities of money” for them).
What you’re trying to do here is quintessential gaslighting.
Even if uploaded to iCloud (such as pictures sent via WhatsApp by default), and above the threshold, they would still be scanned by a second algorithm and subject to human review. So, your "very simple" attack fails on at least three counts.
And again, there is the whole fact that you received the email and there is a log that you received it.
Additionally, the images sent to review are significantly downscaled versions of the original & could easily be made to be ambiguous.
The most difficult challenge in this SWAT story is that Apple has a secret secondary hash that's checked on a collision. That's the part of the SWATting story that feels difficult on a technical level. However, there are also really smart people out there so it wouldn't surprise me if a successful attack strategy is developed at some point given time.
No one is going to be prosecuted on "significantly downscaled ... ambiguous" versions of original fake images with a hash collision that flagged a review and was handed to the FBI accidentally because a "minimum wage" fatigued person passed it on.
I get the counter-arguments, but the hash collision thing is just, sort of... weird? I even get the argument that an innocent hash collision may have your personal and private images reviewed by some other human - and that's weird. But I can't really see it going further (you'll be arrested and sentenced to life in prison from the HASH COLLISIONS!).
It's just using technical terms to scare people who don't understand hashes and collisions and probability, and not really founded on reason.
Which typically will be a court case, or at least questioning by police. This can be quite a destructive event on someone's life. Also, there's no mechanism for whitelisting outlined in the paper, nor can I imagine a mechanism that would work (i.e. now you've got a way to distribute CP by abusing the whitelisting fingerprint mechanism or you only match exact cryptographic hashes which is an expensive CPU operation & doesn't scale as every whitelisted image would have to be in there).
Also, your entire premise is predicated on careful and fair review by authorities. At scale, I've not seen this actually play out. Instead either the police will be underworked & not investigate legitimate cases (too many false positives) or they'll aggressively police all cases to avoid missing any.
(I think) the complaint is: how is that different from just using CSAM images, no collision required?