Apple defends anti-child abuse imagery tech after claims of ‘hash collisions’
vice.com
vice.com
Hash collision in Apple NeuralHash model - https://news.ycombinator.com/item?id=28219068 - Aug 2021 (542 comments)
Convert Apple NeuralHash model for CSAM Detection to ONNX - https://news.ycombinator.com/item?id=28218391 - Aug 2021 (155 comments)
For a company that marketed itself as one of the few digital service providers that consumers could trust, I just don't understand how they acted this way at all.
Either there will be heads rolling at management, or Apple takes a permanent hit to consumer trust.
There's a simple answer to this right? Despite everyone's reaction, Apple genuinely believe this is a novel and unique method to catch CSAM without invading people's privacy. And if you look at it from Apple's point of view that's correct: other major cloud providers catch CSAM content on their platform by inspecting every file uploaded, i.e. total invasion of privacy. Apple found a way to preserve that privacy but still catch the bad people doing very bad things to children.
> other major cloud providers catch CSAM content on their platform by inspecting every file uploaded, i.e. total invasion of privacy.
> Apple found a way to preserve that privacy ...
So scanning for CSAM in a third-party cloud is "total invasion of privacy", while scanning your own personal device is "preserving privacy"?
The third-party clouds can only scan what one explicitly chooses to share with the third-party, while on-device scanning is a short slippery slope away from have its scope significantly expanded (to include non-shared and non-CSAM content).
While local, on-device content access has no such guarantees - it's a question of policy rather than technical infeasibility.
And policies change.
> They are scanning files that are being uploaded. So, yes.
Even considering the situation as it is today, the cloud providers (as well as Apple) are scanning just the files that are being uploaded. So I'm still not sure how one is "total invasion of privacy" and the other "preserving privacy".
> is a short slippery slope away
Obviously people who trust Apple aren't concerned about slippery slopes. What's the point of your post?
There’s this bizarre notion that using end-to-end encryption can absolve you of responsibility, that the authorities will have to accept an answer of “we literally can’t access it”.
That’s just not the case for centralised things: you’re deliberately facilitating some service, government will find you liable for some things in its operation, and if you don’t comply, they’ll fine or shut you down. E2EE doesn’t absolve you from law; law is all about saying you’re not allowed to do things that are physically possible.
(Decentralised things, now they can be banned but not truly stopped because there’s no central party to shut down.)
A state that does not accept it might retaliate against the entity giving that answer or forbid future use of end-to-end encryption without backdoors, but the truth of the answer doesn't depend on anyone's acceptance.
Problem is that the law is self-contradictory and it is up to the judicative institutions to fix it as soon as possible.
You can still do secure backups of your phone without using iCloud, but there isn’t a way for Apple to do end to end encryption of backups transparently like you can with real time communication. The only way end to end encryption of backups works is to require people keep a separate secure key(s) to avoid losing their data, which means a universal implementation has real direct risk for users.
As long as Apple has access to these files the FBI can legally require them to do these searches. From a pure PR perspective they should have communicated what was already going on before releasing this system because people assume something significant changed.
There is no reason the password can't be the encryption key, with backup keys stored with a trusted third party (eg: your credit union or bank) without notation as to what these backup keys are tied to.
Apple has already proven it will concede to China's demands.
They are building the worlds most pervasive surveillance system and when the worlds governments come knocking to use it ... they will throw their hands up and feed you the "Apple complies with all local laws etc.."
I think I'm not the only one who'd rather not have my devices call the cops on me in a country where the cops are already way too violent.
I don’t mind my one drive being scanned for “bad stuff”, I very much mind my personally owned data stores being scanned, with no opt out.
If you turn off iCloud Photos, nothing is scanned.
Microsoft scans everything.
>The system that scans cloud drives for illegal images was created by Microsoft and Dartmouth College and donated to NCMEC. The organization creates signatures of the worst known images of child pornography, approximately 16,000 files at present. These file signatures are given to service providers who then try to match them to user files in order to prevent further distribution of the images themselves, a Microsoft spokesperson told NBC News. (Microsoft implemented image-matching technology in its own services, such as Bing and SkyDrive.)
https://www.nbcnews.com/technolog/your-cloud-drive-really-pr...
According to Apple. and For Now. Patriot Act was only for terrorists. Apple makes concessions for China. Creating this technology, makes it very easy for China to go, "Look at all photos, always". If they only want to scan stuff on iCloud Photos, no worries, just implement on it on their end. This tech does not need to exist in that case.
> Microsoft scans everything.
Everything uploaded to the Cloud. The Cloud bit is fairly important here. That's me, willingly, putting information on their property. Perfectly acceptable and fine for them to ensure that it's nothing unethical.
However, someone else snooping through your drawers looking for something to pin on you is not private, nor the same as checking their own property.
By your definition of "everything", both Microsoft and Apple scan "everything". (Or Apple will, after this new system is rolled out.)
Nobody is arguing about the legality of CSAM itself. It's an issue that is not popular to discuss, and of course being on the wrong side of it results in near-universal, justifiable derision. Stopping its spread is absolutely the right thing to do.
So at the point that companies will always be held liable for storing it, they will have to put up countermeasures of some kind or find themselves sued out of existence. Server-side scanning is one method, and Apple's on-device scanning is another.
There are certainly ways that Apple can go too far with whatever it happens to come up with as its solution to stopping CSAM, but there is still seems to have been a line behind which nobody particularly cares how the detection is implemented and life continues as usual. If Apple had chosen not to cross that line, maybe many of the arguments being made here would never have been brought up at all.
What it takes to accept is the "nothing to hide" mentality: your files are safe to scan (locally) because they can't be known CSAM files. You have to trust the scanner. You allow the scanner to touch your sensitive files because you're not the bad guy, and you want the bad guy be caught (or at least forced off the platform).
And this is, to my mind, the part Apple wasn't very successful at communicating. The whole thing should have started with an educational campaign well ahead of time. The privacy advantage should have been explained again and again: "unlike every other vendor, we won't siphon your files unecrypted for checking; we do everything locally and are unable to compromise your sensitive bits". Getting one's files scanned should have become a badge of honor among the users.
But, for some reason, they tried to do it in a low-key and somehow hasty manner.
You may catch a few perverts looking at the stuff, but I'm not convinced this will lead to catching the producers. How would that happen?
Compare this to the war on drugs. There is huge demand for drugs. Many drugs are highly illegal to possess, and a lot of people get jailed for possessing them. And yet most of us consider the war on drugs to be an abject failure that hasn't done much more than create inequity.
Why should something like CSAM possession be any different? I wish it was different, and these sorts of tactics would reduce this awful problem, but I'm just not convinced it does.
What’s novel about this? The technique and issues with it are fairly obvious to anyone with experience in computer vision. It seems not too different from research from 1993 https://proceedings.neurips.cc/paper/1993/file/288cc0ff02287...
The issues are also well known and encompass the whole subfield of adversarial examples.
Saddened by the privacy-adverse functionality on handsets but Apple seems to be hitting the nail on the head that poor communication principally is driving outrage.
So it's not even accomplishing "keeping it off the servers."
Apple's actions will be effective at making sure every CSAM aficionado will not trust their device. Or, if they're tech savvy, they will at least know not to co-mingle their deepest darkest secrets alongside photos of their mum and last night's dinner.
If you think Apple's very big, very public blow-up is making waves in the hacker/security/libertarian crowd, just imagine how big it's blowing up in the CSAM community right now. I dare say it's probably all they've been talking about for the past two weeks. Unlike places like Hacker News where there's plenty of people desperate to drag Apple through the coals, the CSAM community will be highly motivated to have a precise understanding of what Apple is doing. I dare say that they'll be extremely well informed about exactly what Apple's plans entail and how to work around them.
This is Orwellian doublespeak.
You don't preserver someone's privacy by snooping on their devices.
That is very unlikely. Most likely they compare some hash against a database - just like apple.
So does Apple.
EDIT: some people don’t like that answer, but “inspecting” in this context clearly means “digitally inspecting” (Google does not physically look at every file) and Apple does this with files that are uploaded. They do it in device, but it’s still inspected. That’s the whole point of this controversy, that there’s not much difference to people WHERE apple inspects and on device is actually arguably worse. Your sentence does not in any way distinguish what Apple does from what others do.
They're not going to find predators, all they are going to find is people incompetent enough to have years old (otherwise how would it end up at NCMEC?) CSAM on their iPhones and dumb enough to have iCloud backup turned on. To catch actual predators they would need to run AI scanning on phones for all photos and risk a lot of false positives by parents taking photos of their children on a beach.
Plus a load of people who will inadvertently have stuff that "looks" like CSAM on their devices because some 4chan trolls will inevitably create colliding material and spread it in ads etc. so it ends up downloaded in browser caches.
All of this "let's use technology to counter pedos" is utter, utter crap that is not going to save one single child and will only serve to tear down our freedoms, one law at a time. Want to fight against pedos? Make photos of your airbnb, hotel and motel rooms to help identify abuse recording locations and timeframes, and teach children from an early age about their bodies, sexuality and consent so that they actually have the words to tell you that they are being molested.
Who do you think has 30 or more CSAM images on their phone?
But let's also be real about something else. Think of the Venn diagram of child abusers and people who share CSAM online. Those circles are not identical. Not everybody with CSAM is necessarily a child abuser. Worse, how many child abusers don't share CSAM online? Those can't ever be found by Apple-style invasions of privacy, so one has to wonder if we're being asked to give up significant privacy for a crime fighting strategy that may not even be all that effective.
The cynic in me is also wondering what fraction of people working at places like NCMEC are pedophiles. It'd be a rather convenient job for them. And after all, there's a long history of ostensibly child-serving institutions harboring the worst kind of offenders.
CSAM scanning is sold as beneficial, while in reality it won't do shit while it opens dangerous precedence backdoors!
Apple announces that it is going to start scanning iCloud Photos only, and that their system is set to ignore anything below a threshold of ~30 positives before triggering a human review, and people lose their minds.
This is the difference between putting CSAM on a sign in your front yard (maybe not quite front yard but I can't come up with quite the same physical equivalent to a cloud provider) and keeping it in a password protected vault in your basement. One of those things is protected in the U.S. by laws against unlawful search and seizure. Cloud and on your device are two very different things and consumers are right to be alarmed.
I'll say it again, if you are concerned with this privacy violation, sell your Apple stock and categorically refuse to purchase Apple devices. Also go to https://www.nospyphone.com/ and make your voice heard there.
I honestly can't find the uproar here. Google devices can face match photos offline... so they are applying a neural net (scanning) ON THE DEVICE! How is that not worse than what apple do?
The difference can also be seen from a customer service perspective. One is a feature that lets you sort according to which friends you were with. The other is a feature that puts you in jail. No thanks. Not gonna pay money for that.
Let's say the TSA were to install air-travel-contraband scanners in everyone's homes, but promise only to scan things that are being put into your luggage as you prepare to go to the airport. And let's say that this became a requirement if you want to board a plane.
That's what this feels like. I'm fine with Google scanning through everything in my GMail account, or everything I've uploaded to GDrive, or created in GDocs. That stuff is on their servers, unencrypted, and I explicitly put it there.
But I'm sure as hell not going to let Google install something on my laptop (or phone!) that lets them look at my stuff, even if they pinky-promise that they'll only scan stuff that I intend to upload.
First of all, you have to be able to read it to do the comparison that can increment the counter to 30. So regardless of whether it is or is not encrypted there, they're accessing the unencrypted plaintext to calculate the hash.
And yes, on my device is definitively more private than on someone else's server--just like in my bedside drawer is more private than in an office I rent in a co-working space.
This doesn't matter because Apple can read iCloud data, including iCloud Photos. They hold the encryption keys, and they hand over customers' data for about 150,000 users/accounts a year in response to requests from the government[1].
maybe you mean to say that apple says they won't read it until that threshold has been crossed.
When that scanning gets moved from the cloud to being on your device, a boundary is violated.
When that boundary is violated by a company who makes extreme privacy claims like saying that privacy is a "fundamental human right"[1], yes, people will "lose their minds" over it. This shouldn't be shocking at all.
You would have to have 30 false positives before Apple can see anything, which is unlikely, but the next step is still a human review, since it's not impossible.
Apple can't decrypt the results of the scan until the ~30 image threshold is crossed and a human review is triggered.
Given Google's reluctance to hire humans when a poorly performing algorithm is cheaper, are they turning over every single false positive without a human review?
Facebook, even more so, they are explicitly anti-privacy to the point of being insulting.
Microsoft will happily show you everything they may send when you install Windows, you can sometimes refuse, but not always. They are a bit less explicit than Google, but privacy is rarely on the menu.
As for Amazon, their cloud offers are mostly for businesses, different market, but still, for consumers, they don't really insist on privacy either.
So that if any of these company scan your pictures for child porn, it won't shock anyone, because we know it is what they do.
But Apple claims privacy as a core value, half of their ads are along the lines of "we are not like the others, we respect your privacy, everything on your device stays on your device, etc...", they announce every (often legitimate) privacy feature with great fanfare, etc... So much that people start to believe it. But with that, people realize that Apple is not so different from the others after all, and if they bought an overpriced device based on that promise, I understand why they are pissed off.
https://techcrunch.com/2014/08/06/why-the-gmail-scan-that-le...
You don't consider the contents of your email account or the files you mirror to a cloud drive to be your own private data?
I expect that my ISP tracks and stores my DNS resolutions (if I use their DNS) and has a good understanding of the websites I visit.
I expect that an app that I grant access to my contacts uploads as much data as it can to their servers.
I expect WhatsApp and similar apps to collect and upload meta data of my entire photo library such as GPS info the second I give them access.
Hence, I don’t give access. And hence, it’s a problem if there is no opt-out of local file scanning in the future.
https://protectingchildren.google/intl/en/
> CSAI Match is our proprietary technology, developed by the YouTube team, for combating child sexual abuse imagery (CSAI) in video content online. It was the first technology to use hash-matching to identify known violative videos and allows us to identify this type of violative content amid a high volume of non-violative video content. When a match of violative content is found, it is then flagged to partners to responsibly report in accordance to local laws and regulations. Through YouTube, we make CSAI Match available for free to NGOs and industry partners like Adobe, Reddit, and Tumblr, who use it to counter the spread of online child exploitation videos on their platforms as well.
> We devote significant resources—technology, people, and time—to detecting, deterring, removing, and reporting child sexual exploitation content and behavior. Since 2008, we’ve used “hashing” technology, which creates a unique digital ID for each known child sexual abuse image, to identify copies of images on our services that may exist elsewhere.
Mainland China will probably be the first chip to fall. Can't imagine the Ministry of State Security not actively licking their lips, waiting for this functionality to arrive.
Imagine if Apple had done this on the client side without telling anyone, and later it was discovered. I think things would be a whole worse for Apple in that case.
In part because people didn't know.
And if you were one of innocent people caught by them, you wouldn't want people to know.
Conjecture, admittedly:
1. Apple cannot lose the Chinese market. Huge and more important fastest growing geo for the company.
2. China is self deprecating elements of its own tech sector (aggressive crackdowns on both established companies like tencent and Alibaba as well as individual web sites). They are clearly cleaning house from a surveillance and control perspective. Apple is not immune, but it’s an American behemoth, so open door crackdowns are impossible.
I don’t think Apple’s CSAM push and China’s crackdown are purely coincidental.
Who can argue with stemming child abuse? It’s the type of hot button issue that affords broad acceptance for intrusive tech.
The leap from scanning for abuse to scanning for anti regime content is more like a tiny step.
It seems obvious from afar that the company adamant about refusal to unlock a potential terrorist’s iPhone on privacy principles (with the attendant marketing benefits) would so suddenly force push (and therefore ensure collection massive training data with or without opt-in for Chinas v2.0) such a boldly invasive feature addition.
Turns out vertical integration is both gift and curse (dependent on the whims of the integrator) for on-device privacy and autonomy.
If Apple happens to use similar ToS in China (no idea if true), you bet CCP would be all over this clause.
A worst-case wild guess from the pessimist in me: they had to add this clause in 2019 to appease CCP, and then they got thinking if they could make use of it for the greater good (tm), too. Hopefully it’s very incorrect.
"Apple does not have my permission to use my device to scan my iCloud uploads for CSAM"
and
"This is a slippery slope that could result in Apple enforcing thoughtcrimes"
Neither of these viewpoints are particularly agreeable to the general public in the US, as far as I can determine from my non-tech farming city. Once the fuss in tech dies down, I expect Apple will see a net increase in iCloud adoption — all the fuss we're generating is free advertising for their efforts to stop child porn, and the objections raised are too domain-specific to matter.
It's impossible to say for certain which of your outcomes will occur, but there's definitely two missing from your list. Corrected, it reads:
"Either there will be heads rolling at management, or Apple takes a permanent hit to consumer trust, or Apple sees no effect whatsoever on consumer trust, or Apple sees a permanent boost in consumer trust."
I expect it'll be "no effect", but if I had to pick a second guess, it would be "permanent boost", well offsetting any losses among the tech/free/lib crowd.
"Apple has created a system for detecting CSAM on local devices which has already proven vulnerable to cheap perceptual hash collision attacks. It's now highly inconceivable Apple will be able to deploy this technology as-is without having their users exploited."
In other words it's not just about privacy or thoughtcrimes anymore but should be viewed as actually dangerous to use their devices. I feel a bit dramatic even typing that out but I.. think it's true?
How dramatic is too dramatic? When does something that hasn’t happened to you or anyone you know become a risk you’re willing to sacrifice personal convenience to mitigate? Will you be divesting yourself of all wireless radio hardware? If not, then why would you be worried about users being exploited through a more clumsy and less effective process such as CSAM signature hacking?
The piece of information you’re taking for granted, that few in free/tech/lib are confronting, is the assumption that this process can be exploited at scale to harm millions of people.
So far as I can tell, there will probably be zero or one false positive CSAM matches that pass the known algo, the unknown algo, the human blurred comparison, and the human unblurred comparison — all steps that must occur before law enforcement is invoked to collect digital evidence - in the first year.
How many false positives (to the nearest 10^X) do you think the system will generate in the first year that result in law enforcement actions? Your words suggest that everyone is vulnerable, and there are 10^9 users, so do you believe there will be 10^9 false positives in the first year? Do you think only a thousand people will be affected, so 10^3? How do you judge which is more likely correct?
It is unlikely that this system will generate 10^9 false positives, or else it never would have passed QA. I encourage you to consider how you would personally quantify this risk, and then also look up the quantified risks for killing someone while driving a car or getting struck by lightning while indoors. I don’t know what the actual reality will be, but I don’t think it's a very large X.
If I'm honest what makes me feel bad about this is probably just that- a feeling based on what I consider to be an algorithm designed around the presumption of guilt. It's much the same way I feel about taking my shoes off in line at airport security. It's an act which I've largely come to ignore but which still produces that vague feeling of discomfort that somehow feels like the opposite of security.
That this is occurring on my Apple devices - my favorite devices - is also just depressing.
I mentioned I was thinking of moving from an Android phone to Apple soon, somewhat privacy related.
My friends lectured me on "they're scanning your photos" ... meanwhile they share their google photos albums with me and marvel about how easy they are to search ...
Maybe we (humans) only get outraged based on more specific narratives and not so much the general topics / issues?
I don't know but they didn't seem to notice the conflict.
Psychologically, you'll feel a difference in what you accept between the two, I think
Don't want your photo's scanned, don't sync them to icloud. Seriously! Please include the actual system when discussing this system, not your bogeyman system.
"To help address this, new technology in iOS and iPadOS* will allow Apple to detect known CSAM images stored in iCloud Photos."
To increase privacy - they perform the scan on device prior to upload.
"for now"
Which is the part most people have a problem with -- they say that they are only scanning iCloud uploads now, but it's simple extension of the scanner to scan all files.
I don't care if Apple scans my iCloud uploads on iCloud servers, I don't want them scanning photos on my device.
I'm not sure there's a real difference unless you want to watch your settings all the time. In google land they tend to reset ... and really that happens a lot of places.
I think for most people if you use google, you're in their cloud.
No ... They scan everything that I have released to apple photos that exists on my device.
Same scan - different place.
How did you determine that their intentions contradict their words? Please share the framework for your belief, so that we're able to understand how you arrived at that belief and to evaluate your evidence with an open mind.
(Or, if your claim is unsupported conjecture, please don't misrepresent your opinions and beliefs as facts here at HN.)
This is just my biased view but there _could_ be countless ways to slip the slope without violating the current framing. For example:
Q: What happens when other governments ask Apple to use this for other purposes.
A: We will inform them that we did not build the thing they’re thinking of.[1]
Note that the question could be interpreted as "use this exact implementation" (rather than use the algorithm in general), and the answer does not rule out the possibility of "but we can build the thing for them then" (rather than "we don't have it and we will never have it"). The reader is free to interprete the conversation as they see it.
[1] https://daringfireball.net/2021/08/apple_child_safety_initia... citing [2]
[2] https://www.nytimes.com/2021/08/05/technology/apple-iphones-...
Definitely would appreciate a link to anything substantial indicating that this was a bunch of eng in over their heads.
I’m not buying the “engineers were left unbridled” argument, I think there just have been a level of obliviousness in much wider a part of the organization for something like this to happen.
Maybe because they underestimated people's ignorance.
I think they saw (and still do see) it as a better, more privacy preserving technique than what everyone else is doing.
Their hubris is in not seeing or thinking that they will be able to stand up to all kinds of abuses of this system that its mere existence will invite; their hubris is also in thinking that they will be able to perfectly and without mistakes manage and overview a system making accusations so heinous that even the mere act of accusing destroy people's lives and livelihoods.
Since the announcement, I can think of a dozen ways Apple could be easily forced into scanning all the contents of your device by assembling features they’ve already shipped. Yet they haven’t. At some point, people need to produce evidence that Apple cannot hold the line they’ve said they will.
What we have now is Apple, with its "strong privacy" record, normalizing this. If it succeeds, it would be that much easier for the governments to tackle other stuff onto it. Or, say, lower the threshold needed to submit images for review. I can easily picture some senator ranting about how unacceptable it is that somebody with only 20 CSAM photos won't be flagged, and won't somebody please think of the children?
And yes, if it comes to that, Apple definitely cannot hold the line. After all, they already didn't hold it on encrypted cloud storage - and that wasn't even legally forced on them, merely "not recommended".
Because privacy stance is mostly PR to differentiate from Google. And while there're invalid reasons to get users data, there're also valid ones (at least from legal requirement point of view - let's not get into weeds about personal freedom here and if the laws and its implementations need to be changed).
Their PR was just writing the checks they cannot cash without going on a war with governments.
I assumed they pivoted to focus on privacy, but clearly it was just a marketing department innovation rather than a core value (as their marketing department claimed).
If your threat model includes being the target of someone who will plant child pornography on your phone, you are already fucked. And no, Apple isn’t suddenly going to scan Chinese iPhones for Winnie the Pooh memes. They don’t have to. China already has the 50 cent party to do that for them, on WeChat.
Basically everything everyone seems to think is just around the corner has already been possible for years.
Most people won't have any idea about the meaning of "hashes" and "databases". Not everyone is trying to actively fight the system and shit on everything, most people just want to live happily with their friends and family, they won't care that Apple scans their devices.
> "Either there will be heads rolling at management, or Apple takes a permanent hit to consumer trust."
Oh god ! How wasn't all of this obvious to the top Apple management, but so obvious to epistasis! Damn, thanks man for correcting and leading Apple to the right track !
What I have seen is a selling point for apple products.
I'd encourage folks to get out of the HN bubble on this - talk to a wife, a family especially those with kids.
Why stop there? Get out of the HN bubble on the patriot act, instead ask your neighbor’s wife her thoughts on it. Get out of the HN bubble on immigration, go ask a stereotypical boomer conservative about it.
I think my sarcasm already made it overtly obvious but, this is horrible advice you are giving and the fact that you don’t seem to be aware that pedophilia and terrorism are the two most classic “this gives us an excuse to exert totalitarian control” topics betrays your own ignorance (or, worse, you are aware and just don’t care).
The only people who are bothered are people claiming this is going to be misused by authoritarian governments.
There are atleast 2-3 further checks to account for this.
Salvador Dali could do something similar by hand in 1973 in Gala Contemplating the Mediterranean Sea [1]
The mere accusal itself of possessing CSAM can be life ruining if it gets to that stage. More importantly, a collision will effectively allow warrantless searches, at least of the collided images.
“Cryptographic representations of images”. That’s not the case though right? These are “neuralhashes” afaik which are nowhere close to cryptographic hashes but rather locality sensitive hashes which is a fancy speak for “the closer two images look like, the more similar the hash”.
Vice and others keeps calling the cryptographic. Am I missing something here?
I know that that "think of the children" is a meme, but I think this illustrates the point for apple. If you have a croped or modified image of CSA the system will identify it. As long as your image is different enough from CSA, you are safe.
The point here is that Apple is specifically looking for matches against known CSA material.
If someone can demonstrate that a legal NSFW image (eg. regular old-fashioned pornography), can be collided with a legal, completely 100% SFW image then I'll be concerned.
But until then, this looks like a reasonable and supportable way for finding CSAM in real time.
Look at this other collision: https://twitter.com/SarahJamieLewis/status/14282060881181491... An attacker can send you an innocent looking picture that embbed some CSA material and you get swatted the next day.
Consider if the honeypot images (manipulated to match CSAM hashes) are terrorist recruitment material for example.
Bonus points if you match poses, coloration, background, etc.
1) These are more share similar visual features than crypto hashes.
2) HN posters have been claiming that apple reviewing flagged photos is a felony -> because HN commentators are claiming flagged photos are somehow "known" CASM - this is also likely totally false. The images may not be CASM and the idea that a moderation queue results in felony charges is near ridiculous.
3) This illustrates why apple's approach here (manual review after 30 images or so flagged) is not unreasonable. The push to say that this review is unnecessary is totally misguided.
4) They use words like "hash collision" for something that is not a hash. In fact, different devices will calculate DIFFERENT hashes for the SAME image at times.
One request I have - before folks cite legal opinions - those opinions should have the name of a lawyer on them. Not this "I talked to a lawyer" because we have no idea if you described things accurately.
Not going to happen. Lawyers in the US have issues with offering unsolicited advice, and other problems with issuing advice into states where they are not admitted. So likely none of the US lawyers (and the great many more law students) here will ever put their real name to a comment.
Lawyers I know would politely decline that.
So your own firm may cover some costs if you have something to say. If you found someone to pay for you to do an analysis or offer your thoughts - you'd be in heaven!
But it's the law, it's fuzzy at best, much like your HR department. It's only after a court decision has been reached on your particular issue that it's anywhere near "settled" case law, and even that's up for possible change tomorrow.
At least HN should flag these and get these taken down. Over and over the legal analysis is either trash or it's clear the article author didn't understand something (so how can lawyer give good advice?).
These conversations become so uninteresting when people take these extreme type positions. Apple's brand is destroyed - apple is committing child porn felonies.
I would have rather just had a link to the apple technical paper and a discussion personally vs the over the top random article feed with all sorts of misunderstandings.
And in contract law there are LOTS of legal articles online - with folks name on them! They are useful! I read them and enjoy them. Can we ask for that here where it matters maybe more?
Articles are not legal advice. They are opinions on the law applicable generally, rather than fact-based advice to specific clients. Saying whether apple is doing something illegal or not in this case, with a lawyer's name stamped on that opinion, is very different.
Agree, and I think this is backed up by real world experience. Has Facebook or anyone working on their behalf ever been charged for possession of CSAM? I guarantee they've seen some. Probably a lot, in fact. That's why we have recurring discussions about the workers and the compensation they get (or not) for the really horrid work they are tasked with.
Overall, we're advised to take about these steps there. First off, report it. Second, remove all access for the customer, terminate the accounts, lock them out asap. Third, prevent access to the content without touching it. For example, if it sits on a file system and a web server could serve it, blacklist URLs on a loadbalancer. Fourth, if necessary, begin archiving and securing evidence. But if possible in any way, disable content deletion mechanisms and wait for legal advice, or the law enforcement to tell you how to gather the data.
But overall, you're not immediately guilty for someone abusing your service, and no one is instantly guilty for detecting someone is abusing your service.
> In 2020, FotoForensics received 931,466 pictures and submitted 523 reports to NCMEC; that's 0.056%. During the same year, Facebook submitted 20,307,216 reports to NCMEC
https://www.hackerfactor.com/blog/index.php?/archives/929-On....
Do you really think Apple's brand has been "destroyed" over this?
They use “private set intersection” (https://en.wikipedia.org/wiki/Private_set_intersection) to compute a value that itself doesn’t say whether an image is in the forbidden list, yet when combined with sufficiently many other such values can be used to do that.
They also encrypt the “NeuralHash and a visual derivative” on iCloud in such a way that Apple can only decrypt that if they got sufficiently many matching images (using https://en.wikipedia.org/wiki/Secret_sharing)
(For details and, possibly, corrections on my interpretation, see Apple’s technical summary at https://www.apple.com/child-safety/pdf/CSAM_Detection_Techni... and https://www.apple.com/child-safety/pdf/Apple_PSI_System_Secu...)
People who understand tech well enough to recognize hashes like MD5 and SHA don't dive deep enough to understand that this is something completely different.
I even suspect this is deliberate from Apple's side when announcing and talking about these changes - making people wrongly believe that only exact matches will trigger, except possibly in extremely rare cases and under concious attacks.
They could have called it "fingerprint" or something but deliberately went with a technical term that even confuses technical people who know well enough what a hash usually means.
Vice is falling victim to this misunderstanding stemming from the conflation of "hash".
It's a huge, huge, huge distinction.
Edit: this is apparently not true as demonstrated by researchers.
> Microsoft says that the "PhotoDNA hash is not reversible". That's not true. PhotoDNA hashes can be projected into a 26x26 grayscale image that is only a little blurry. 26x26 is larger than most desktop icons; it's enough detail to recognize people and objects. Reversing a PhotoDNA hash is no more complicated than solving a 26x26 Sudoku puzzle; a task well-suited for computers.
https://www.hackerfactor.com/blog/index.php?/archives/929-On...
https://twitter.com/fayfiftynine/status/1427899951120490497?...
Given neuralhash is a hash of a hash, I imagine they’re running photodna and not some custom solution which would require Apple ingesting and hasing all of the images themselves using another custom perceptual hash system.
> . Instead of scanning images in the cloud, the system performs on-device matching using a database of known CSAM image hashes provided by NCMEC and other child-safety organizations. Apple further transforms this database into an unreadable set of hashes, which is securely stored on users’ devices.
https://www.apple.com/child-safety/pdf/CSAM_Detection_Techni...
Apple isn't using a "similar image, similar hash" system. They're using a "similar image, same hash" system.
[1]: https://www.apple.com/child-safety/pdf/CSAM_Detection_Techni...
Perceptual hashes are not cryptographic hashes. Perceptual hashing systems do compare hashes using a distance metric like the Hamming distance.
If two images have similar hashes, then they look kind of similar to one another. That's the point of perceptual hashing.
You know, let me put it this way. Yiu know that one really weird family member I'm pretty sure everyone either has or is?
Guess what? They're a neural net too.
This is what Apple is asking you to trust.
IGNORE THIS: I think that's the parent comment's point. These are definitely not cryptographic hashes, since they—by design and necessity—need to mirror hash similarity to the perceptual similarity of the input images.
In order for a collision to get through to the human checkers, the same image would have to fool both networks independently:
Unclear how hard this would actually be in practice (if I were going to attempt it, the first thing I'd try is to evolve a colliding image with something like CLIP+VQGAN) but certainly harder than finding a collision alone.
Swing and a miss. Not in the CSAM dataset. Take two images. Encode alternating pixels. Decode to get the original image back. Convert to different encodings print to PDF or Postscript. Encode as base64 representations of the image file...
Who are we trying to fool here? This is kiddie stuff.
This is more about trying to implant scanning capabilities on client devices.
Did we think they didn't already have that ability?
It's a direct trade-off and the error tolerance of any such filter is the only thing that makes it useful so we can basically stop arguing about the depths of implementation details or how high the collision rate of the hashing algorithm is etc. If this thing is supposed to catch anyone it needs to be magnitudes more lenient than any of those minor faults.
Or better yet, they can just not store their stuff on an iPhone. While meanwhile, millions of innocent people are have their photos scanned and risking being reported as a pedophile.
1. It requires a perfect 1:1 match (their documentation says this is not the case); 2. Or it has some freedom in detecting a match, probably including a match with a certain percentage.
If it's the former, it's completely useless. A watermark or a randomly chosen pixel with a slightly different hue and the hash would be completely different.
So, it's not #1. It's going to be #2. And that's where it becomes dangerous. The government of the USA is going to look for child predators. The government of Saudi Arabia is going to track down known memes shared by atheists, and they will be put to death; heresy is a capital offence over there. And China will probably do their best to track down Uyghurs so they can make the process of elimination even easier.
It's not like Apple hasn't given in to dictatorships in the past. This tech is absolutely going to kill people.
>"This independent hash is chosen to reject the unlikely possibility that the match threshold was exceeded due to non-CSAM images that were adversarially perturbed to cause false NeuralHash matches against the on-device encrypted CSAM database," …
We are just one stupid terrorist attack from full surveillance of everybody.
Neural hash this: Its about Trust. Its about Privacy. Its about Boundaries between me and corporations/governments/etc.
If you are Apple, even though EARN IT failed... you know where Washington's heart lies. Is CSAM scanning a "better alternative", a concession, an appeasement, a lesser evil, in the hope this prevents EARN IT from coming back?
Also, many people forgot about the Lawful Access to Encrypted Data Act of 2020, or LAED, which would unilaterally banned E2E encryption in entirety and required that all devices featuring encryption must be unlockable by the manufacturer. That also was on the table.
If you’re trying to frame this as “we need to prevent Congress from ever passing something like the EARN IT act”, I agree. Apple and other tech companies already lobby Congress. Why aren’t they lobbying for encryption?
It's clear that EARN IT could literally be revived any day if Apple didn't do something to say "we don't need it because we've already satisfied your requirements."
Alternatively, "Apple has shown that it's possible to do without undue hardship, so we should make everyone else do it too".
They were going to legally mandate that everything be scanned through methods less private than the ones Apple has developed here, through EARN IT and potentially LAED (which would have banned E2E in all circumstances and any device that could not be unlocked by the manufacturer). While that crisis was temporarily averted, the risk of it coming back was and is very real.
Apple decided to get ahead of it with a better solution, even though that solution is still bad. It's a lesser evil to prevent the return of something worse.
That tech is not being deployed on iMessage which is the only e2ee(ish) service from Apple (with Keychain) and is what those legislative attempts are usually targeting. One could argue it would have made sense (technically) there though, sure.
Was it a reason to release it preventively, on something unrelated, to be in the good graces of legislators ? I'm not sure it's a good calculation, and it doesn't cover other platforms like Signal and Telegram that would still be seen as a problem by those legislators and require them to legislate anyway.
This would only make sense if Apple intends to expand their CSAM detection and reporting system to detect and report those other things, as well.
Also, there is another reason why there is the CSAM Detecting and Reporting system. With Apple CSAM Scan, that big "excuse" Congress was planning to use through EARN IT to ban E2E is diffused, meaning now Apple has the potential to add E2E to their iCloud service before Congress can figure out a different excuse.
There would need to be end-device scanning for arbitrary objects, including full text search for strings including 'Taiwan', 'Tiananmen Square', '09 F9', and so forth to even begin looking at e2e encryption of your items in the cloud.
At which point… what's the point?
It did not do this. The bill was essentially asking for search warrants to become a part of the protocol. If you're only solution to allowing for search warrants to work is to stop encrypting data I feel you are intentionally ignoring other options to make this seem worse than it is.
Its amazing since this would have decimated the American Tech sector in many unknown ways.
In some countries even discussing the application of certain numbers is unlawful.
But you shouldn’t get jailed for protected speech and you shouldn’t get jailed for preserving your privacy (via encryption or otherwise.) As cynical as people may get, this is one thing that we have to agree on if we want to live in a free society.
And above all, most certainly, we shouldn’t allow being jailed over encryption to become codified as law, and if it does, we certainly must fight it and not become complacent.
Apathy over politics, especially these days, is understandable with the flood of terrible news and highly divisive topics, but we shouldn’t let the fight for privacy become a victim to apathy. (And yes, I realize big tech surveillance creep is a fear, but IMO we’re starting to get into more direct worst cases now.)
Not in any democratic country.
How can you prove that I'm using encryption? How you even define what encryption is to make it possible to have a law that bans it? If me and you we put down a code to communicate, is that encryption and thus banned?
Also you wouldn't need to ban encryption completely. There are cases where encryption is required by law (e.g. digital signature on documents) and others where the encryption is required for things to work (e.g. credit cards, mobile phones, etc).
You have to put down a list of cases where you can use encryption and others where you can't. That is nearly impossible.
If you are IG Farben, you know where Berlin's heart lies...
Does Apple's solution only stop people from uploading illegal files to Apple's servers or does it stop them from uploading the files to any server.
If Apple intends to control the operation of a computer purchased from Apple after the owner begins using it, does Apple have a duty to report illegal files found on that computer and stop them from being shared (anywhere, not just through Apple's datacenters).
To me, this is why there is a serious distinction between a company detecting and policing what files are stored on their computers (i.e., how other companies approach this problem) and a company detecting and policing what files someone else's computer is storing and can transfer over the internet (in this case, unless I am mistaken, only to Apple's computers).
Mind you, I am not familiar with the details of exactly how Apple's solution works nor the applicable criminal laws so these questions might be irrelevant. However I was thinking that if Apple really wanted to prevent the trafficking of ostensibly illegal files then wouldn't Apple seek to prevent their transfer not only to Apple's computers but to any computer (and also report them to the proper authorities). What duty does Apple have if they can "see into the owner's computer" and they detect illegal activity. If Apple is in remote control of the computer, e.g., they can detect the presence/absence of files remotely and allow or disallow full user control through the OS, then does Apple have a duty to take action.
Only applies to iCloud Photos uploads, but the photos are still uploaded: when there's a match, the photo and a 'ticket' are uploaded and Apple's servers (after the servers themselves verify the match[0]) send the image to human reviewers to verify the CSAM before submitting it to police as evidence.
0: https://twitter.com/fayfiftynine/status/1427899951120490497 and https://www.apple.com/child-safety/pdf/CSAM_Detection_Techni...
Perhaps they're trying to prevent being in "posession" of illegal images by having them on their own servers, rather than preventing copying to arbitrary destinations.
Duty? No, that's the secondary question. The primary question is whether they have the right.
All of the slippery slope arguments have already been possible for nearly a decade now with the cloud.
Can someone illustrate something wrong with this that’s not already possible today.
Fundamentally unless you audited the client and the server yourself either (client or backend) scanning is possible, and therefore this “problem”.
Where’s the issue?
That doesn't make it ok. The fact that you and my mom have normalized bad behavior doesn't make it any less offensive to my civil liberties.
> Am I the only one who finds no issue with this?
If you only think about the first order effects. This is great, we catch a bunch of kid sex pedos. The second order effects are much more dire. "slippery slope" sure... but false arrests, lives ruined based on mere investigations revealing nothing, expanding the role of government peering in to our personal lives, resulting suicides, corporations launching these programs - then defending them - then expanding them due to government pressure is *guaranteed*, additional PR and propaganda from corporations and government for further invasion in our personal lives due to the marketed/claimed success of these programs. The FBI has been putting massive pressure on Apple for years, due to the dominance and security measures of iOS.
Say what you want about the death penalty, but many many innocent people have dead in a truly horrific way, with some actually being tortured while being executed. That is a perfect example of second order effects on something most of us without any further information (ending murderous villains is a good thing) would agree on. So many Death Row inmates have been exonerated and vindicated.
edit: https://en.wikipedia.org/wiki/List_of_exonerated_death_row_i...
Not only has it been possible for a decade, it’s been happening for a decade. Every major social media company already scans for and reports child pornography to the feds. Facebook submits millions of reports per year.
Scanning publicly accessible material makes a lot of sense and is within the boundaries of section 230, tech companies are responsible for policing removing illegal or threatening material. For private or encrypted information, section 230 does not apply.
The purposes these services were designed for matters. iCloud is meant to be private, facebook would prefer everything you post is public as it’s built in to benefit their business model. Facebook should scan, iCloud shouldn’t. You are purposely obfuscating that detail.
Apple is introducing a reverse 'Little Snitch' where instead of the app warning you what apps are doing on the network, the OS is scanning your photos. Introducing a 5th columnist into a device that you've bought and paid for is a huge philosophical jump from Apple's previous stances on privacy, where they'd gone as far as fighting the Feds about trying to break into terrorist's iPhones.
The reason people don't like this, as opposed to, for example, Dropbox scanning your synced files on their servers, is that a compute tool you ostensibly own is now turned completely against you. Today, that is for CSAM, tomorrow, what else?
> people don’t care about that because you’re literally, voluntarily, giving Dropbox your files.
Correct and I agree. I don’t upload my most personal photos to Dropbox for this very specific reason. In fact I stopped using Dropbox when Condi Rice joined the board because she lacks good sense and doesn’t respect civil rights. See ‘The Patriot Act’ and ‘the Invasion of Afghanistan’. It was easy to stop using Dropbox because the alternatives were vast. Apple has me very purposely locked in to this scheme to where the alternatives are a huge transition and compromise on privacy no matter where I turn.
Not to mention the meaning of ownership is completely unrelated to ability to view the designs of a given object. I think I own the fan currently blowing air at me without ever having seen a schematic for its controls circuitry just fine and everyone for all of history has pretty much felt the same.
The point is it's hard to hide your evil plans in daylight. Either you trust Apple or your don't. Same with Microsoft's telemetry, they wrote the whole OS, if they were evil they have 10000 easier ways to do it.
All Apple has to do is lie, you'll never know.
The issue is that Apple previously was not intruding into their user's privacy (at least publicly), but now they are.
It sounds like your argument is that Apple could have been doing this all along and just not telling us. I find that unlikely mainly because they've marketed themselves as a privacy-focused company up until now.
Apple reserves the right at all times to determine whether Content is appropriate and in compliance with this Agreement, and may screen, move, refuse, modify and/or remove Content at any time, without prior notice and in its sole discretion, if such Content is found to be in violation of this Agreement or is otherwise objectionable.
You must not have been paying attention for the last 20 years.
Why is that the alternative? How about everything is encrypted and nothing is scanned.
If you don’t want to be scanned you can turn it off. I honestly don’t see the issue. It seems the only thing people can say are hypothetical situations here.
How about neither? Just let people have their privacy. Some will misuse it. Thats life.
What is also possible: No scanning on your device and encrypted cloud storage. E.g. borg + rsync.net, mega, proton drive.
What I'm basically getting at is: are the files scanned after the user has expressed the intention of uploading them? That's what I understood. Am I wrong? Are the files scanned the moment they appear on your device, regardless of you iCloud status (even if you have disabled iCloud somehow)?
Edit: typo
1. Get a pornographic picture involving young though legal actors and actresses.
2. Encode a nonce into the image. Hash it checking for CSAM collisions. If you've found a collision go on to the next step, if not update the nonce and try again.
3. You now have an image that, to visual inspection will appear plausibly like CSAM, and to automated detection will appear like CSAM. Though, presumably, it is not illegal for you to have this image as it is, in fact, legal pornography. You can now text this to anyone with an iPhone who will be referred by Apple to law enforcement.
So at this point we have an image that computers think is CSAM and people think is CSAM, and when held up next to the original verified horrific image everyone agrees is the same image. At this point, someone is going to ask, rightly so, where that came from.
In order to generate this attack, you have had to go out and procure, deliberately, known CSAM. Ignoring that it would be easier just to send that to the target, rather than hiring talent to recreate the pose of a specific piece of child porn (or 30 pieces to trigger the reporting levels), the most likely person by orders of magnitude to be prosecuted in this scenario is the attacker.
Define "get anywhere". Why won't you get raided by the police and have all your devices seized first?
If your 30 or so hash matching images matched their corresponding known CSAM then that goes on to the police and then they knock on your door.
I am under the impression that Apple's scheme allows them only to verify the output of the matching algorithm (the "safety vouchers"), and not the image content itself. So in the hypothetical situation described in the thread, it won't be possible for Apple to detect the false positive, and they could pass on the report to NCMEC.
I fear that this will lead to a law enforcement raid without any actual human verification of the offending image itself.
If I'm wrong, I welcome citations which demonstrate the opposite.
You're already attempting to frame someone for a crime, might as well commit another crime while you're at it
Room for improvement in the headline.
Isn't that where we already with things like Article 17 of the EU's Copyright Directive?
When did they make NeuralHash public?
So they are already running a generic version of this system since iOS 14.3?
I think the claim here is that you won't have access to the source images, and therefore generating collisions will be more difficult. But, if you do have access to the source images, this has been shown to be trivial. This of course doesn't stop nations states generating images that cause hash collisions, in fact they would be incentivized to do so.
I would also add that Apple are behind the curve, attempts to crack the hashing algorithm more efficiently are still ongoing: https://github.com/AsuharietYgvar/AppleNeuralHash2ONNX/issue...
> [..] not the final implementation [..]
Why on earth would you invite people to come and test your algorithm and then say "sure, you broke it, but it's not the real one". This kind of defeats the point and seems like some bait and switch bullshit. I suspect this is some retroactive cope from management realising they can't deploy this version and whatever they do deploy needs to be heavily modified.
> If Apple finds they are CSAM, it will report the user to law enforcement.
One statistic I want to know is: How many people already trigger this report function in the wild? Surely currently is the largest number of positives they will ever have - if it turns out to be 0% - Apple should just scrap it.
> Apple also said that after a user passes the 30 match threshold, a second non-public algorithm that runs on Apple's servers will check the results.
So to avoid reporting, simply block Apple servers? Also, security by obscurity is not security - the algorithm supposedly being private just means that its not properly tested and Apple is not held to account.
> "Apple actually designed this system so the hash function doesn't need to remain secret, as the only thing you can do with 'non-CSAM that hashes as CSAM' is annoy Apple's response team with some garbage images until they implement a filter to eliminate those garbage false positives in their analysis pipeline," Nicholas Weaver, senior researcher at the International Computer Science Institute at UC Berkeley, told Motherboard in an online chat.
No. A report could be considered 'reasonable doubt' for law enforcement to do a full search. Imagine trying to explain to a judge why your iPhone shouldn't be searched because of a false-positive CSAM hash collision because of a malicious website you visited or a text message you received.
Sounds pretty desperate.
As if any normal user is going to upload a photo to iCloud that is a collision.
The fact that such images are possible to generate means nothing by itself.
Also, they would need to accidentally have 30 of them.
Also, a human would have to not be able to tell the difference.
If Apple is training a neural network to detect this kind of imagery, I would imagine there to be thousands, if not millions of child pornography images on Apple's servers that are being used by their own engineers to train this system
NCMEC generate the hashes using their CSAM corpus.
I'm not saying this as a fan of either Cook or their anti-CSAM measures; I'm neither, and if anyone is ever wrongfully arrested because Apple's system made a mistake, Cook may well wind up in disgrace depending on how much blame he can/can't shuffle off to subordinates. I don't think we're there yet, though.
Just saying that money hides problems.