Google is testing its new ad targeting tech in millions of browsers
eff.org
eff.org
The proposal rests on the assumption that people in “sensitive categories” will visit specific “sensitive” websites, and that people who aren’t in those groups will not visit said sites. But behavior correlates with demographics in unintuitive ways. It's highly likely that certain demographics are going to visit a different subset of the web than other demographics are, and that such behavior will not be captured by Google’s “sensitive sites” framing. For example, people with depression may exhibit similar browsing behaviors, but not necessarily via something as explicit and direct as, for example, visiting “depression.org.” Meanwhile, tracking companies are well-equipped to gather traffic from millions of users, link it to data about demographics or behavior, and decode which cohorts are linked to which sensitive traits. Google’s website-based system, as proposed, has no way of stopping that.
The way I interpret this is that, based on your browsing history in Chrome (or any browser that implements this kind of functionality) you are placed into a number of categories (or, if one reverses the metaphor, a number of descriptive tags are attached to you). Google is aiming to ensure that certain categories/tags that might be considered sensitive (mental state, physical illnesses, etc.) will be blocked.(To be clear, this is my interpretation of what they are stating, not an assertion of fact)
The EFF is arguing that this isn't really that straightforward, as sensitive details can still be inferred from non-sensitive details.
What I'm curious about is, who is doing all the ID generation, categorization, and data centralization? Or is Chrome just going to calculate everything itself, then send the data to sites that ask for it?
A FLoC is a single ID per browser and it's not human readable, the algorithm to determine this will run in the browser so your data will never leave it. There's no real way for someone to "reverse engineer" the ID, it's just a random set of bits not "people-who-like-shoes".
Last, and I may be wrong here given this announcement, FLoC isn't part of Chrome until Chrome 91 in june, at least I think this is what the Chrome team said Tuesday at the W3C meeting so I'm not sure what this refers to, but it shouldn't be the browser side.
Those parasites pay for the vast majority of the internet you likely use and benefit from.
Yes, but I don't want them to. Of course, I don't have a choice. And the people that keep supporting them keep helping them work their fingers deeper, slowly touching and corrupting everything that I might still find enjoyable there. That's what makes them parasites. They take what's good and they make it worse.
Pay and subscribe to all the sites you read or video sites you watch and let's see how many sites are left that you read. Incidentally I subscribe to a lot of online news sites and video sites to avoid ads where they bother me or to support the journalism, but not everyone can afford this, especially smaller publishers.
If you have a new business idea that represents this third way... What are you doing here? Build a business and become a billionaire.
The internet of the mid 90s had a lot of stuff for free. Usenet was great for certain groups of people. The internet of the mid 90s also lacked a ton of products that cost a ton of money to develop and maintain.
People are frightened now of relying on Google to store their photos forever. Imagine storing your photos with a passion project operated by just some person in their spare time. We see people already getting pissed off when OSS developers stop maintaining projects. It'd be that problem times 1000.
This person claimed ignorance of the other party once and in context it's a very reasonable conclusion. There's nothing rude about that, except maybe when you find the state of ignorance bad in and of itself. That doesn't really make sense, as all humans are born ignorant. That's just how things are and that's what teaching is for. Which is what the post being discussed attempted to do.
You could get some users together and fund this by a group.
In particular they aren't happy also about the 3rd party cookie comparison because it just makes no sense, not sure why the PR focuses on that given that this is minimum 3-4 orders of magnitude less precise than 3rd party cookies, targeting a FLoC is going to be like targeting a zip code, hardly THAT privacy invasive.
No, it's going to be like targeting by zip code when you already can target by other means as well (fingerprinting, IP address). So it's going to make fingerprinting easier.
Each browser is assigned a single floc ID, which is (for now) a 50-bit hash that is not human readable on its own. But the floc ID is intended to capture something meaningful about browsing behavior. Two users with similar browsing history are likely to have similar, if not the same, floc IDs. Floc IDs which are close together (hamming distance-wise) also reflect similar browsing histories. (http://matpalm.com/resemblance/simhash/)
It will be totally possible to gather data about what kind of people belong to a particular floc -- provided that you operate a large website or ad network. It will probably not be possible for normal people.
There's a risk that advertisers will learn that a particular floc ID corresponds to a "sensitive category" of users. That's what google is trying to prevent.
Google plans to audit floc IDs for correlation with visits to sensitive websites. In order to do so, it's assigning a "sensitivity" label to certain websites, and then applying that label to each user who visits those websites. Then it will run analyses to see whether users with a particular floc ID are more likely to have a particular label. Floc IDs that correlate too strongly with a particluar label will be thrown out.
The passage is attempting to explain why this does not prevent adtech from learning other meaningful things about flocs -- for example, that users in a specific floc are disproportionately female, or Muslim, or suffering from depression.
As for whether FLoC is in chrome, see the origin trial page here: https://developer.chrome.com/origintrials/#/view_trial/21392...
Regarding the trial, yeah I know that origin thing is up, and the minutes for yesterday's W3C IWABG meeting aren't going to be out until next week, but the team working on this stuff at Google said that the trial up there is kind of fake and the code isn't merged in Chrome yet and won't be until 91, at some point I have to trust the engineers working on this stuff more than the rest.
[edit]: To be clear, I don't mean that there's no experiment going on, I just mean that it's not the browser running it, that the removal of 3rd party cookies as a way to opt-out likely is due to the fact that they may be relying on server-side computation to do this experiment rather than the browser, I wouldn't see why the browser would care about cookies to track your history, it's already stored in the browser.
Yes, and given enough data this:
> so long as you can't tell sensitive categories
will not be true.
> I have to trust the engineers
The engineers that work for a giant corporation whose entire reason for existing is mining and collating as much data as possible on you? Why would you trust them at all?
For example, insurance claims. I hope Google isn't tracking enough information to test for correlations between FLoC cohorts and insurance claims, but insurers will have access to that info when customers visit their site to pay bills.
Perhaps they'll find they can more accurately anticipate customer's claims from their FLoC cohort, and use that to price new policies, thus indirectly pricing policies based on browsing history.
The insurance company has the data, and it doesn't even need to know any of the details, it just needs to realize that certain cluster numbers correlate with future claims.
Perhaps Google will think of some of these correlations, and flag articles about becoming parents as sensitive, but how many unexpected correlations might be overlooked?
I’m not sure how the insurance company gets that data though in your example.
AIUI, they can get customers' FLoC cohort when customers visit their site to pay bills (or do anything else), and see prospective customers' FLoC cohort when issuing quotes online.
Who knows what they'll be able to learn from that? I certainly don't, but I expect them to try, and that worries me.
Isn't bundling users into buckets good for individual privacy? In theory if the bucket is too small you are identifiable, but my understanding is that the entire premise of this approach is to ensure that is not the case.
None of those issues apply to paid journalism.
And even without clickbait stuff people search for the news that satisfies them and their value system.
All of this happens on data voluntarily shared with the platforms people complain about like Facebook, and there's no end to this in sight.
Advertising in this forum is at the same time useless and barely works (cue the 50% advertising is wasted, don't know which one, and other similar sentences) and the most powerful tool for mind control ever existed.
Any revenue model based on the number of views (like advertising) is fundamentally flawed and will suffer from the exact same problem.
The trend towards segmenting people into ever-smaller classes to serve them more targeted ads is another cause of, not a solution to polarization.
Most traditional adtech, used to targeting attributes, is quite puzzled at how to use this at the moment given they'll all need to invest in machine learning because these IDs aren't interpretable.
The specifications of this stuff are all in the open.
That's the key though, isn't it? While I can see that the intentions may be good, this isn't being released into a vacuum, but rather to a world where all of that other tracking infrastructure already exists.
One day I receive a request that says you are group 12345, in 1 hour you are 12346 instead because the algorithm recalculated your FLoC... Not sure what anyone can tell about that without knowing the group ids meanings, which likely not even google will know since this is going to be a clustering algorithm anyway.
As long as browser fingerprinting exists it will be possible to track the FLoC cohorts a person belongs to over time and learn much more about them than a single FLoC cohort reveals.
But AFAIK, FLoC is only replacing third-party cookies; browser fingerprinting will still be possible.
Effective fingerprinting makes everything impossible. Even if there were no 3rd party cookies and no FLoC, ad-tech can still use fingerprinting to bucket you based on your browsing history. This isn't a criticism of FLoC, this is a criticism of unwanted fingerprinting.
Sites that know who you are (e.g. because you log in) gain even more info, if they previously only knew about your visits to that site.
The fingerprinting would fail to recover the same code at least once a day unless you always visit the exact same sites every day and never visit a new one.
fingerprinting needs a stable state from the browser to work, not something that changes arbitrarily.
But beyond simply fingerprinting, it still is providing some tracking. I've seen a lot of "anonymized" data turn out to be not so anonymized in practice, or reversible to a large degree.
Such lack of imagination... The recorded data build behavioral profiles of the whole globe on multiple dimensions.
I absolutely hate it and can't find anything about it online, other than 2-3 people on twitter complaining about it as well. No idea how they're even doing this or how I can disable it.
Its bad but it's a fact. How many decades of outrage is it gonna take for you to accept it.
You can opt-out of this (for now) in Chrome by disabling third-party cookies.
You can also simply use another browser such as Firefox.
I think Firefox has improved significantly in recent years as well, but I haven't used it in a while.
The flock system sounds like the Chinese social credit score. I'm wondering what things will be conditioned on FlockID. There are going to be elite flocks and worthless flocks. "Sorry, our services are available only to flock-3453 and flock-2234. Losers like flock-23232 need not come."
your youtube is slow ? , oh honey Just click on these few websites to change your flock id to one of a more advertisement friendly one.
Looks like they found one the most computationally cheapest way to discriminate against users on the internet.
One identifier to discriminate them all :)
Not sure if it would suffice to just overwrite the document.interestCohort(); function and have it report something trollish.
Since a cohort is rather small could a botnet be sufficiently large, to create its own private cohorts and mess up a lot of add deliveries?
It seems like if each user only gets a single floc grouping, that this is more private than a system like FB, where any given user could be part of thousands of different targetable "interest groups." Am I missing something? On FB, for instance, I could be targeted for liking Infinite Jest. And separately for like Mountain Biking. And separately for living in Pennsylvania. It seems like Google is doing a lot to obscure the user information into a data black box. Maybe I don't get the idea.
This effectively solves the question of optout as i can choose to use default value so that i am indistinguishable from thousands other people This also clearly allows for some targeted ads that user does actually care about. I don't mind seeing ads for technology, but all those "You wouldn't believe this!! 11", and "Look, penis!!" are just insult to humanity.
I know it's still ads. But i have an impression it's ao much better solution
Any different advertising platform will always be inferior by definition.
Edit: PS: I when I'm searching for a product, placed results (which reasonably match my intent or possible related needs) aren't ads. Those are helpful suggestions curated by payment and hopefully regulated match.
E.g. basic Gmail will be free, but if you want something like Undo Send, etc. then you'll have to pay a monthly fee for it.
And these fees from the various sites will add up quickly.
This isn't at all clear, mostly because every "ad" professional I've talked to IRL is as allergic to real data as some people are to peanuts. What you've quoted is a sales pitch.
Here's a review of all the relevant literature on the topic: https://docs.google.com/presentation/d/1PKHVtO6hgwBJS1vafLyv...
Where a giant group is compared to an extremely tiny group and the differences between them are used to make sweeping generalizations about everyone. When this is the quality to expect, it's just good practice to be skeptical.
Edit: Even if this weren't the case, the enormous spread of estimates when it comes to revenue impact really says it all.
'No ads' would be preferable.