[1] https://9to5google.com/2020/02/06/google-chrome-x-client-dat...
[1] https://9to5google.com/2020/02/06/google-chrome-x-client-dat...
See this post I made when X-Client-Header was introduced.
> But they claim [1] this X-Client-Data header is used for experimenting with Chrome, not for tracking.
They claim a lot of things. Sometimes they even modify their claims years after they first made them. Even if they were making 100% innocent claims now, they are not guaranteeing[2] they won't change how they use the data in the future.
> But you're claiming they're using it for tracking, so you're claiming they're lying...
Not at all - as I point out in [1], Google is saying they are tracking people with that header, but the claim is obfuscated by some blatant doublespeak and the hope that you only consider the X-Client-Header (or any other field with >0 bits of fingerprintable entropy) in isolation. That is never true, <i>which Google also admits</i> when e.g. they casually mention they can deduce a HTTP request's country of origin.
Asking if the X-Client-Data is "used for tracking" is the wrong question. They are tracking people using the combined set of all of the tiny pieces of data they are able to exfiltrate from you. Any specific piece of data isn't important; if that header was missing or corrupt, it's just a small amount of noise added to the already noisy correlations they do to fingerprint you to infer whatever they are interested in about your pattern-of-life.
No they don't. Nothing in your comment sounds anything like Google claiming to be tracking people. And as mentioned elsewhere in this thread, they explicitly claim to not be tracking individuals.
>> "The information included in this header reflects the variations, or new feature trials, in which an installation of Chrome is currently enrolled. [...] it is not used to identify or track individual users."
I believe this statement is true. They are not using the X-Client-Data HTTP Header to track "individual users". As they state in their "Privacy Whitepaper"[3], a "low entropy variation" is
>> randomized based on a number from 0 to 7999 (13 bits) that's randomly generated by each Chrome installation on the first run.
This per-installation 13 bit id number is sent in HTTP requests to certain Google domains:
>> [...] a subset of low entropy variations are included in network requests sent to Google. [...] These are transmitted using the "X-Client-Data" HTTP header. [...] This header is used to evaluate the effect [of the variation (presumably?)] on Google servers [...]
Google is explicitly saying they are tracking the new 13 bit id number. This particular id number is somewhat low granularity, due to the limited bit length, which - as Google correctly claims in their statement to the press - is too small to "identify or track individual users". Regardless, they are still claiming they track a 13 bit identifier tied to "each Chrome installation".
I never said the X-Client-Data header was enough to track individuals. An IP address doesn't uniquely identify individuals either. Google's use of language is hoping you stop there. The header is too small to be identifying, so it doesn't matter! This framing is only true if you limit your questioning to considering the "low entropy variation" number in isolation. If that was the only number Google was able to track, it indeed wouldn't be concerning. However, that number is probably transported to Google over the internet in an IP Protocol packet, meaning they are at a minimum also receiving either a 32 bit (IPv4) or 128 bit (IPv6) identifier in the Source Address field of each packet's IP Protocol header.
Google doesn't need to use the X-Client-Data header to "identify or track individual users". They can simply use it to disambiguate different Chrome installs that share the same public IP address. This usage of the number isn't identifying users; it's only identifying the different Chrome installs e.g. behind a typical household stateful NAT router. Both the X-Client-Data header and the IP address are both not unique personally identifying IDs. However, the tuple {X-Client-Data, IPv4 Address} is probably unique for most people. In the rare instance where it isn't unique, one of the related tuples like {X-Client-Data, IPv4 Address, User-Agent} will be.
The doublespeak is pretending the header "will not contain any personally identifiable information" when the stated purpose of header is to create a new tracking identifier that accomplishes the same thing as a personally identifying identifier when you combine it with the other data that Google already tracks (such as the 24 bits of "anonymized" IP address (they zero the LSB) that they store with each GA record.
[3] https://www.google.com/chrome/privacy/whitepaper.html#variat...
> when the stated purpose of header is to create a new tracking identifier that accomplishes the same thing as a personally identifying identifier when you combine it with the other data that Google already tracks
This is not stated anywhere except by you. The stated purpose of the identifier is to track analytics around chrome experiments, and only that.
The tuple (IPv4 Address, User-Agent) is already unique for almost all, so why go to all the effort?
Google are totally lying here.
If you never go to doubleclick yourself, chrome won't ever send data to it. It's not like a sneaky background thing. It's extra data attached to requests you were already making.
To a first approximation, _nobody_ "goes to doubleclick themselves".
At the same time, back in 2016 a study at Princeton found almost 50% of all sites on the web had Doubleclick on them (this is separate to the 70% of sites running Google Analytics - and I'd bet there's approximately zero sites that serve doubleclick ads/trackers but not google analytics ones, so there's even less way to spin this as being a necessary way for google to "understand experiments" by whitelisting the doubleclick domain...).
I never "go to doubleclick". My browser "sneakily in the background goes to double click" while I browse about half the sites on the internet.
"It's extra data attached to requests you were already making." is technically true, and gaslighting at it's most brazen.
If you asked your mom how many times she made a went to doubleclick today, what would she say? What would the actual answer be if we wanted to use the tortured terminology of "requests she was already making to doubleclick" from your apologia about your employer up there?
You took me to task for calling you a stooge and that being "against HN policy" elsewhere in this discussion...(Well, I said "stooge", you accused me of saying "shill", but whatever.) I guess I apologise for using the term "stooge" for someone who's told us they work at google and are telling lies about how Chrome is sending unexpected tracking data to Doubleclick. But it seems very much the right term.
How about instead I say that your statement "If you never go to doubleclick yourself, chrome won't ever send data to it." is a brazen lie, wrapped up in a weasel-wordy disingenuous interpretation of what 99.99% of people would clearly understand "go to doubleclick yourself" to mean?
I'll leave this argument now, with a quote for you and all googlers:
"It is difficult to get a man to understand something, when his salary depends upon his not understanding it!" -- Upton Sinclair
(Source for my 50% number: https://www.technologyreview.com/2016/05/18/160139/largest-s... )
> At the same time, back in 2016 a study at Princeton found almost 50% of all sites on the web had Doubleclick tracking on the
Means that many people are going to doubleclick, via it's ads existing on other sites. If the extent of your concern is that I said "going to" instead of "makes requests to", valid and I apologise for not being precise in my use of language.
However, the HN guidelines also ask that you respond to the strongest possible interpretation of what someone is saying. So please respond to what it is clear I meant, and not the straw man you feel compelled to attack.
And since that bit of rhetorical drama seems to at this point be the core of your concern, I don't know that theres anything of substance for me to address here, just more personal attacks.
> My browser "sneakily in the background goes to double click" while I browse about half the sites on the internet.
To be clear, this is false. Your browser isn't doing anything sneaky here. Perhaps you can argue that individual websites are being sneaky by including 3p advertising. That's a valid concern. But a browser "loading the HTML of the page you direct it to" isn't being sneaky or nefarious.
We all know the answer. Zero. Whatever your sneaky-browser-behaviour-depending employer wants you to say in public, her answer will be zero. Same as 99.99% of the world.
I think the behavior on display is pretty sneaky, undermines privacy and users aren't really informed about Google doing "research" on them.
But the question here is what part of this is sneaky? Is including a weird tracking header sneaky? Perhaps. Is making requests to doubleclick sneaky? In the context of those requests being made as part of a page load? Not on the part of Google, which is my point.
The original complaint was that the header was being sent
> That Google need to send "experimental headers" to a hardcoded domain for an advertising company they bought a decade or so back - because of course the results of web browser experiments should go to an advertising company
And the reason why is simple: that's the domain people are already making requests to.
"The X-Client-Data header is used to help Chrome test new features before rolling them out to all users. The information included in this header reflects the variations, or new feature trials, in which an installation of Chrome is currently enrolled. This information helps us measure server-side metrics for large groups of installations; it is not used to identify or track individual users."
Okay, so they're saying it is not currently being used to track individual users.
This means they can:
- Track individual users in the future.
- Track devices at any time, for differentiation of individual profiles.
- Track installs at any time.
- Derive information about the individual using this data (for example how often they update their browser, how often they use that device etc)
All of those things are extremely valuable from a data standpoint.
Unless Google explicitly stated that it was not being used to track, differentiate or profile users or devices and will not be used for that purpose in the future, it's extremely suspicious.
Most ad networks collect bulk information from browsers to fingerprint devices and users. Google has a step up on the competitor networks because they own the browser.
Do not forget Google was sued just this month for tracking users while they were using incognito mode.
If you want to say you think they're tracking people because of reasons X/Y/Z, or that what they're doing looks suspicious, that's a lot better of an argument. But to claim they are with no evidence of it actually happening is really pushing it.
Every single word of that statement was carefully crafted and constructed. Knowing that, why is it so ambiguous?
> Every single word of that statement was carefully crafted and constructed.
I also don't believe this to be true (their statement seemed plain and clear enough to my eyes), but even if it were, it doesn't affect what I said above. You need actual evidence, not the mere possibility of mathematical loophole.
They may not even be explicitly lying, because the statement is so ambiguous. When you're reading PR/legal speak then every single word matters.
Sorry, but you're not going to.
> They may not even be explicitly lying, because the statement is so ambiguous. When you're reading PR/legal speak then every single word matters.
Then say it's ambiguous, instead of saying the opposite is true, is all I'm saying. You misinform people that way.
Google lost the benefit of doubt years ago. This is Google 2020 - all they do is "track users":
https://www.reuters.com/article/us-alphabet-google-privacy-l...
https://www.compliancejunction.com/google-loses-appeal-of-e5...
> I might as well claim you're a burglar because there's nothing to indicate you're not one.
Google have been caught burgling houses repeatedly, and have been found in your house with burglary tools claiming "we're just doing, ummm, _browser experiments!!!_" You saying there's "no evidence of it actually happening " isn't useful.