Facebook is reportedly trying to analyze encrypted data without deciphering it
engadget.com
engadget.com
https://en.m.wikipedia.org/wiki/Homomorphic_encryption
I'm not sure how the author jumped to the conclusions they did, but they show a severe misunderstanding of the applications of homomorphic encryption
https://github.com/tf-encrypted/tf-encrypted
Homomorphic means the algebra is the same on both sides of the encryption, so this would be one of the types of encryption where this is possible.
In reading through the readme, its data the user encrypts (which is a goal of homomorphic encryption), not learning on data of someone else that is encrypted.
Assuming data us properly encrypted, it is indistinguishable from noise, so I don't see how someone can learn on someone else's encrypted data.
It's all about homologies, understand what that means to understand why this is possible while still having good security and doing encryption correctly.
With regard to FHE, I'm familiar with the concept in abstract, but this article and the comments made me wonder how traditional definitions of security work in a homomorphic world.
Specifically, do we assume under FHE that the computation provider has no a-priori information beyond basic structure of the values contained? (i.e. do you need to exchange some key to the party who should be able to do computation, even if this key can't decrypt the ciphertext to plaintext?)
Assuming this isn't the case, do we retain classical definitions of security (the ciphertext leaks no information about plaintext; large collections of ciphertexts similarly reveal no information about plaintexts, or the similarity or difference between any two plaintexts) under FHE?
It seems if this holds true, FHE mathematically can't do this, as it would require revealing information from a ciphertext that means it is no longer effectively encrypted, hence it isn't homomorphic.
However, in practice, you reveal some information about the data if you do computations in it. One notional example is if you see a bunch of XORs in data. You may not know what the data is, but if you see a bunch of XORs, you can make an educated guess that some sort of cipher operation occurs (as XOR is used heavily in it).
They would have to get information from other observations, e.g. the fact that you send (encrypted but nonetheless observable) messages to certain people at certain times and frequencies, telling them something about the apparent importance of that relationship.
1. Length correlation of messages - a block cipher will be padded to the nearest block, a stream cipher could (absent padding) give you an exact length while preserving confidentiality. Length lets you infer a very high level category of content (for example, that's likely a short video, that's likely a long video, that's likely a gif, that's probably a still image).
2. Homomorphic encryption as the article suggests, or something simpler - client-side signalling of targeting parameters. Let's say I have the IAB list of 1000 ad targeting categories and I make a list of keywords for these. You could scan messages in plaintext at client-side, via the app that holds the keys. For a given message (the provider knows who you're communicating with, unless the platform is designed more like Signal, to avoid metadata). You can assign a few potential IAB keyword categories to a given conversation. Those can be encrypted into a little block "addressed" to the service provider, and appended to the message sent from the device. Good luck spotting this without static analysis of the app.
If you could do this, you could add ads to a given conversation view based on IAB "categories" assigned to that conversation or group chat. You could also infer categories for a given user if there's a clustering of categories present throughout a series of conversations. And that's all just based on client side keyword scans.
That's technically extracting meaning/value from encrypted data, without attempting to go homomorphic. And while FHE is cool technology, I'm not sure it makes sense commercially to pursue that if you could just quietly sneak client side targeting in and get away with it due to user apathy and regulatory stagnation.
In essence if you use FHE and FB can gain any inference about information contained therein, it's no longer E2EE as the ciphertext conveys information about the plaintext absent the key. If you wish to make a leaky cipher, you could just client-side add the metadata for advertising or whatever other purpose, as either way you need to "backdoor" the client to achieve this information leakage.
As you rightly say, the avalanche property of a good cipher, combined with basic principles of cipher mode initialisation and non-reuse of use-once parameters, and use of message authentication schemes that don't leak a MAC of the plaintext at ciphertext level (like MAC-then-encrypt does), means all an observer should see in common between two identical messages is the length. The rest should be entirely avalanched (with exception perhaps of a sequence header in the transport layer depending on the protocol, but even that could be encrypted)