1) This still lets you have personalized models, just trained on more than 1 user, thats fine at google's scale anyway
2) Their competitors (FB, AMZN) dont have the edge compute (Android) to do this, and to a lesser degree don't have the ML stack (however Android implements this at the API level will be very Tensorflow focused)
3) Now google can push for privacy regulations that prevent FB and AMZN from storing your raw data
4) Profit
That said theres nothing stopping FB doing federated learning within their app on mobile, I just don't think they have the privacy background to bother.
The cynic in me sees this as a great play by the Googz to cut the Amazon-adtech venture in the bud, and to establish and maintain dominance over the adtech business, and advertising in general.
It's not Android scale but Amazon has sold 100 million Alexas.
I started reading it looking at it from a cynical view, but ended it with a "hmm... this could actually work," especially after reading about Secure Aggregation. That's badass.
Of course it depends on the application... For example, I might not want my keyboard autocomplete learning from others, but I might want my self driving car to do it.
For applications where we want a common model, I see no other way to do it but this. The idea companies collect massive stockpiles of data forever is infeasible.
For example back when recaptcha was introduced, trolls tried to transcribe everything as the n-word (since you only had to get one of the two words correct).
These cases can obviously be noticed and fixed but it's harder now that the training data is opaque right?
Cynical view is that it's only used when Google doesn't want your individual data. This will help muddy the waters in discussions about privacy.
I.e. you can't necessarily make the decision locally if it's not local problem, and ad-targeting usually isn't a local problem.
In fact, if this is seriously gonna be used as a "better privacy" argument (as some people seem to be already doing in this very thread), I'm calling it a PR victory for the "bad guys".
First off, if you were worried about what Android was sending to Google, there's no reason to believe it's going to stop. In fact, I believe that the biggest problem always was the (carefully cultivated) confusion about what data is being sent: you have a hundred of menus to "opt-out" of something and it's not even exactly clear if it changes anything. In fact, we know for a fact, that when "opting out" in some cases more data is being sent. And even if we are not entirely happy about it, most of us still allow this to happen, because there's nothing we can directly do to prevent it and everybody says "well, I do need a smartphone after all, right?" (yeah-yeah, somebody doesn't, but we are not talking about the weird minority here)
And all of it happens when it's relatively straightforward to see what data is being sent, because of minimal aggregation on the device. And still even somewhat technically-minded people don't really know what Google (Facebook, Amazon, whatever) really knows about them.
Second, what really is "federated learning"? Well, let's imagine no humans speak Chinese, but there is this program (owned by Google), that does. And speaking Chinese is how it actually operates internally, when deciding to show this or that ad to you, or sending a ballistic missile to your location. So, in order for it to learn, we normally were sending English sentences, which were translated to Chinese server-side. Federated learning is when they are translated to Chinese client-side (which might be considered "lossy" conversion, but to what degree is not really specified), and then sent to Google to be aggregated.
So, yeah, no raw data has been sent, but the central Chinese-speaking machine still somehow knows it all. What exactly it knows, depends on what we really meant by "translating into Chinese" in our metaphor. But effectively, we just offloaded some processor work to the client side, which, as I said, seems really cool to me from the technical perspective, but there's no way it automatically protects us from anything.
Third is basically 1 + 2: we didn't know what is being sent when it was all raw-data, we will know even less, when it's client-side aggregated in some unintelligible-for-the-humans way. And it scares me even more, because if it allows some PR victories for the Google&Friends, then sky is the limit for what more surveillance can be done this way. I mean, if it would be publicly known that Android sends all the sound and all the image from your mic & camera to Google, I think (I hope!) that people would seriously oppose to that. But if it's not the real images, but just some matrix of weights, learnt from them — it might be less clear if anybody has to object to that. And I think they absolutely have to! Because if we don't make any very restrictive assumptions about what we mean by "learning" in this very specific case, then the only thing that matters is that the "central brain" still saw all these images, it just isn't known what exactly it "remembered".
After all, we, humans, also don't store all the pictures we've seen in our brains: it doesn't make you much happier if I saw your transaction history (or whatever else you don't want me to know), because it was never the picture I was after, but only the "aggregated info".
You might be tempted to object, that it can learn nothing about your private life, since it isn't even known which device sent what data. But I think this is silly, because you forget the most important thing, which is to ask: what ever are "you"? The statement in question depends on definition of that, because if "you" means "your IP address", it might be true. But I never really was afraid Google is learning something about what is sent from this IP, because I don't believe the give a fuck.
It is much less apparent they don't learn anything about "John Doe" simply given the fact that the learning is "federated". But then again, this doesn't really bother me much, because I don't believe they care.
I think, much more probable definitions of "you" that might concern them, are a lot more dangerous, because they are a lot more "real" than your made-up (even if at birth) name which you probably share with 100 more people around the world anyway. Like, for example, "the guy, who every day makes a trip from Baker street 12 to Sesame street 30" (that's "you") and that he really likes Coke. And this is just the most simplistic one, there are infinite tuples of parameters that would define a specific person in much more meaningful way, than the IP or a name.
But, actually, I don't really believe they would try to identify you like that either. Well, they might, I just don't really believe they care about "you" in a sense that might be meaningful to you. What some entity like Google is likely to mean by "you" is probably several organisms, so you might feel like whatever they learn about "you" is definitely not private data. But I don't think this is less scary, quite the opposite, facts like "boys 13-15 y.o. that listen Tokyo Hotel and drink Sprite are likely to try heroine if recommended to watch RocknRolla" are much more powerful and useful (this is obviously a made up example, which you may replace by anything seemingly less dramatic, like "person, who buys A, B and C will also want to buy D"). Anyway, what really makes up a model that can control the financial markets and mood of the people isn't about your petty definition of "you", but a much colder, more meaningful one.
And when I'm worried Google learns something about "me", this is what I'm really worried about. And the fact that such a definition of "private data" wouldn't hold up in court because it "isn't even about a specific person" makes it only so much worse.
I don't even feel necessarily comfortable with the completely anonymous usages of ML on my data. Like the mentioned "next word prediction". Language models we've seen by now don't really understand anything about the text, and surely nothing about who you are. Yet they are uncannily good in "understanding" the context somehow. It really doesn't know anything about the world in the strict sense of the word, but given your sentence starts with "Tensorflow" it is still able to understand that something about "neural networks" and "machine learning" would be a good way to continue.
So if it learns on some very unique stories on a very unique subject you were telling someone in the WhatsApp, a model, trained on this data, actually might tell someone else the story vaguely resembling what you just said given the right context. Even though it didn't try to learn anything specifically about you, or even gather any data from you in a non-anonymous way.
Of course, I don't mean to say this is actually likely to happen with how it's likely to be used, I'm just saying that to illustrate the possibility.
If you find it concerning that it's possible to predict behavior based on demographics, I don't know what to tell you. Do you think that psychology studies are an invasion of privacy?
How about we lay off the conspiracy theories every time any new paper is published.
It's like your responding to the Bitcoin paper by saying "proof of work" doesn't mean a thing because "hashing" isn't "money".
There are papers published, critique those specifically, instead of relying on a handwavy respond to a comic.
I wouldn't be critiquing your response if you had something a little less handwavey to say. When people point out flaws in protocols, I like to see specific exploit examples, like the kind you'd put into a Spectre/Meltdown Advisory.
If you said "out of order speculative execution in CPUs might eventually allow exploits", I may even have vaguely agreed, but without a concrete criticism, it's more of an "uneasy feeling" you have.
I can make loose arguments too. Everything you do in this word leaks entropy. Your information is entangled with other people, leaving a wake behind you from the moment you're born. There's a gazillion side channels hanging off of you. So learning information about you is pretty much a given. The question is, is it relevant or important information? There's a huge difference between "learning something", "learning something sensitive about me (that I share with a large number of other people)" and "learning something sensitive about me that's individually traceable to me"
Most people won't care if a Federated Learning model, using data from your phone, learns that people who stay up late, and search for Coke, also end up with diabetes, anymore than a double blind study learns about the risks of smoking and lung cancer -- the doctors have learned information about you (you're a smoker, and you have/don't have lung cancer), but they haven't learned that you, krick, are a smoker.
The whole point of differential privacy, federated learning, and other techniques, is that aggregate statistics, and aggregate models can be learned without any personally traceable information.
Now, you could argue that somewhere, deep within the logical depth of the weights of a DNN is some kind of personal information that could be deanonymized, but this is like a claim that you found a weakness in a hash function -- until you show it, it's just a claim, and mathematics and security research is full of wrong claims on both sides.
Differential privacy is exactly meant for this, in fact. Differential privacy adds a certain amount of randomly-generated noise to client inputs. The result is that, statistically speaking, it’s impossible to tell the difference between a model with your data in it and a model without your data in it.
Arguably the reason the comic doesn’t mention differential privacy is that it’s neither new nor invented at Google. Or maybe just because it’s not technically part of federated learning. But the “federated learning at scale” paper Google put out mentions it, and says they have implemented DP techniques.