>In today's Big Tech antitrust hearing in front of a Congressional subcommittee, representative Val Demings questioned Google CEO Sundar Pichai about the company's merger with DoubleClick.
Specifically, the way Google combined data from the advertising company -- bought in 2007 -- with Google's own data. Founder Sergey Brin had told Congress it would not combine the personal information, but the company quietly did so in 2016 anyway.
https://www.engadget.com/google-antitrust-hearing-doubleclic...
>Google’s privacy policies as of March 1, 2012 established that no combination between DoubleClick’s advertising data and Google’s personally-identifiable information would take place without the prior consent of its users, but an updated version of those policies subtly allowed the integration of both databases regardless of prior consent of its users
https://www.promarket.org/2020/08/21/why-we-should-be-carefu...
how does this make the following OK?
> No Pebble user ever consented to their health data being sold to Google, but all of this is legal.
Entirely possible I misread the spirit behind the letter of what GP wrote.
I think this is part of why Google has earned itself a reputation for killing things: if you require everything to use the same set of internal systems, and can't just leave old projects limping along on a fork of a deprecated system, the cost of keeping old things around is higher.
(I worked at Google until June)
The thing is that these processes don't actually work, because in the end someone is going to do something dumb anyways and you really can't stop them. The same way your processes aren't going to stop an outage, someone is going to correlate data they shouldn't and somehow lack the tact to go "hey maybe I should not be doing this". Often various pressures will make it seem like something they should do. And when they do it they will cost the company an incalculable amount of user trust, to say nothing of the irreversibility of a privacy incident. Someone will do a writeup of it after it is discovered by the press but at that point the damage is done.
(FWIW, I'm not even including the stuff that goes on with executive approval, where they make decisions on how likely they think they are going to be discovered and sued for large amounts of money. Of course this would bypass any policy…but it's usually kept under wraps for obvious reasons.)
This has the downside that the number of organisations that need to approve any launch ends up being borderline overwhelming. This is a big part of why it takes so long, if at all, to get Google products outside the US. Another implication is needing the expensive security org and culture.
You could argue they had no choice, but I could argue that they had the choice to allow search history to be end to end encrypted and/or anonymous. That is if their business model did not rely on selling plaintext PII.
https://www.nbcnews.com/news/us-news/police-google-reverse-k...
As a Googler that has access to some Fitbit stuff, I have to sign a form every 30 days that indicates I understand this.
ML is where things tend to be a little bit more grey overall since being able to look at data is very useful for development, so some things are scrubbed for PII, but then accessible in some form. But for things like GMail or Photos, I would assume nobody (including ML engineers) can read your data as these are basically impossible to sanitize.
Some products have systems train ML models without engineers seeing the data, e.g. spam filtering, even when the underlying data is considered sensitive.
The fairly recent case of people's private conversations being shipped out to basically unvetted contractors for labeling and analysis (and subsequently leaked) should serve as sufficient evidence that "shit happens," and if private conversations without having even initiated an interaction with your Google devices are being tossed around and leaked, forgive me if I don't believe that when producing the tagging, timeline, and album features in Google Photos there wasn't some underpaid, unwatched contractor snooping through my photos without my permission.
I reported the bug. Knowing how security works technically, I added to the bug the words "I'm happy with whoever works on this to take a look at the gallery, here's a world-readable sharing link". A couple rounds of bug comments later, I have been asked to sign a legally binding consent form allowing an engineer to look at the gallery. Then somehow they decided I need to sign a different form to satisfy whatever other legal spirit needed appeasing. Only then someone finally looked at the bundle of photos. They figured out whatever was triggering the bug. They generated a gallery reproducing the bug with generic sample images. Whoever worked on the bug and adding a regression test worked off that synthetic gallery instead.
And that, moreover, I get none of the profits from these transactions, and have no control over whom it's sold to and under what terms.
But somehow this is ok because within Google the common employee doesn’t have access, but is rather incentivized to make the vacuum or selling more efficient?
When people talk about 'security' at Amazon/Microsoft/Google/Apple, they're describing audit performance and ISO compliance. From a business perspective, that's really the only thing that matters. Everything else is played pretty fast-and-loose. Both Apple and Microsoft have been caught backdooring their infrastructure for foreign governments, if anyone out there still has faith in these companies then I hope they stop taking life advice from VC heads.
All data collection practices are not equal.
>Google Tracks 39 Types of Private Data, the Highest Among Big Tech Companies
Google takes the cake when it comes to tracking most of your data. This should not surprise, given that their entire business model relies on data.
Twitter and Facebook both save more information than they need to. However, with Facebook, most of the data they store is information users enter.
Apple is in a league above Amazon in protecting user privacy. It is the most privacy-conscious firm out there. Apple only stores the information that is necessary to maintain users’ accounts.
https://stockapps.com/blog/google-tracks-39-types-of-private...
Apple could have folded, like even (old) Google did, and refused to comply with the CCP but shareholders would have been too upset seeing the price go down.
Successful publicly traded companies will almost always do the most profitable thing legally possible, no matter how evil.
But that’s not the real danger of Google having this data. The danger is in ML. Google’s entire business model is predicated upon using information about you to change your behavior, and to sell on the market predictions of your future behavior.
Google is one of the few companies known to have caught, fired, and officially publicly named an employee who did something like this:
https://techcrunch.com/2010/09/14/google-engineer-spying-fir...
Most tech companies about which people don't routinely raise this type of concern have far weaker security controls against (and detection systems for) this threat model than Google does.
That said:
> But that’s not the real danger of Google having this data. The danger is in ML. Google’s entire business model is predicated upon using information about you to change your behavior, and to sell on the market predictions of your future behavior.
I'm of two minds of this. I'm not thrilled about how much data Google combines and unnecessarily insists on collecting in order to allow things like the Google Assistant and Google Maps to provide full functionality. At the same time, many of Google's assistance and search services are better than their competitors exactly for this reason. I primarily wish they were more transparent in this area with fewer dark patterns and more user control, with forcing users to pick between excessive data sharing and inadequate access to Google services.
Disclosure: I have worked for Google in the past, but not since early 2015. I certainly am not speaking for them here.
> I'm of two minds of this.
I'm of two minds on this.
> with forcing users to pick between
without forcing users to pick between
Not disagreeing with your broader point, but that specific employee was not caught by Google but was reported externally by parents of the minors after abusing access for months. I suspect the incident you linked predates - and was the impetus for - many controls that were subsequently added
Speaking speculatively: I've never worked for Google
(I left Google in June.)
That all mentioned here clearly means google doesn't have data privacy at the core of their priorities and this won't change unless forced by fines/regulation, just like banks. Slightly disappointed when reading this, but I guess I shouldn't have expected more.
What's wrong with granting engineers access to PII-less logs[1] by default? How does that compromise privacy?
I'm willing to bet your bank differed on the following ways from Google in absolute numbers and per-engineer:
* handled significantly less requests per second - at least 2 orders of magnitude - therefore lower log volumes
* shipped less changes to prod per unit time, so fewer problems to investigate
1. No IP client address, raw session id or username
For using it internally: health data is widely considered toxic. As in "I'm not touching that thing with a barge pole" toxic. I would personally be pretty surprised if Google ever started monetising it.