Hash collisions are 100% impossible to avoid. You are mapping an infinite set (the set of all possible images) to a finite set (a fixed length number).
Cryptographic hashes are designed so that collisions are hard to construct at will. But this is not a cryptographic hash at all and I wouldn't be surprised that constructing an image that matches a given hash is easy.
>Hash collisions are 100% impossible to avoid. You are mapping an infinite set (the set of all possible images) to a finite set (a fixed length number).
If you want to be needlessly pedantic I guess. But for 99.999999% of usage hash collisions don't exist in practice.
>Cryptographic hashes are designed so that collisions are hard to construct at will. But this is not a cryptographic hash at all and I wouldn't be surprised that constructing an image that matches a given hash is easy.
You got a cite to back this up? Because they claim otherwise.
Oh right, that very soothing. The very device I spent $1000 on has 0.0001% of ruining my life by causing a no-knock raid due to a false positive. They should put this stuff on their ads, makes me wanna buy more Apple products.
I'd much rather sell all my current Apple devices and permanently switch to linux than do this.
If you're correct, and the chance of a hash collision is 0.0001% then for each hash in the database, that's 200,000 collisions.
Assuming the database has 1,000 hashes [2] and there's no overlap in collisions for any given person, that 200 million peoples' lives ruined.
[0] rough guess based on: https://www.statista.com/statistics/276306/global-apple-ipho... the exact number might be off but I think the order of magnitude is about right.
[1] My gf has taken at least 3 photos every day for the last 5 years at least (so well over 5000 unique photos) - 200 is very conservative.
[2] Again, very conservative. This is a collection of all known child porn images. I wouldn't be surprised if the number is actually two orders of magnitude higher.
[edit] replied to the wrong comment. Bugger. Hopefully contributed to the conversation anyway.
>less than a one in one trillion chance per year of incorrectly flagging a given account
If we gave every single human being an iPhone we'd expect an incorrect flag every 160 years or so.
And I guess we'll find out.
The horrible thing is, of course, that this whole process will probably be automated and there'll be no recourse to a human with any power to work things out properly. If there are thousands of cases, they'll have to take to Twitter to shine some light on it. Except that in this case they're self-identlfying as suspected child abusers. How many people are going to risk the negatives that go with that? Will we ever find out how many people were actually false positives?