Protecting Against Browser-Language Fingerprinting
brave.com
brave.com
Every time I see tech being proposed for privacy instead of legislation I wonder how the topic is kept so vague. There are a handful of companies that can track you across multiple websites, so any real solution has to start by enumerating and addressing those companies.
It feels like were always discussing about curing "diseases" without explicitly saying that malaria, TB, etc are the targets.
I think it's more like hundreds than handfuls. But they're all connecting your behavior across sites using the same few techniques:
* Explicit methods: cookies, link decoration, and other browser-supported ways of adding entropy. Browsers are working on removing these, but if they move too aggressively here then adtech just moves to:
* Fingerprinting: using existing browser entropy. Generally worse than explicit methods because the user doesn't have control (ex: shared fingerprint between successive private browsing sessions). Browsers are also working on reducing this, see the article, but it's very hard because the number of techniques is large and they generally use features users/sites depend on.
* Timing attacks (pretty sure no one is doing this commercially yet)
You might be interested in https://github.com/michaelkleber/privacy-model
(Disclosure: I used to work on ads at Google)
Could you be a bit more explicit here? I'd be curious how these entities coordinate themselves, since I don't have hundreds of cookies set by any website, so there must be few networks to correlate this data (FB, Google, some IAB groups, who else?). Going after those networks would seem like the obvious next step then.
Thanks for the link. It seems like an honorable direction, but it's a bit nebulous on how browsers would be incentivized to implement with good intentions.
Are you sure? These are third-party cookies, and it's not easy to get a full list. One way to do it is to go to a major publisher (NYT, CNN, etc) with devtools open and networking enabled. Filter to third party requests and look for ones sending cookies. Trying this on the NYT front page I saw 3p requests with cookies to amazon-adsystem.com, doubleclick.net, prebid.media.net, rubiconproject.com, adnxs.com, 3lift.com, openx.net, google.com, scorecardresearch.com, casalemedia.com, pubmatic.com, bluekai.com, adsrvr.org, bing.com, twitter.com, everesttech.net, criteo.com, dotomi.com, bidswitch.net, mfadsrvr.com, agkn.com, pswec.com, adtdp.com, demdex.net, bidr.io, adition.com, brand-display.com, intentiq.com, w55c.net, pippio.com, rlcdn.com, and adsymptotic.com before I got bored and stopped counting. Some of these might not be for personalized advertising, but most of them look like it.
> browsers would be incentivized to implement with good intentions.
Browsers compete on privacy, and what they do is open source. So while their incentives aren't perfect, external groups (and competing browsers!) can help keep them honest by paying attention and calling attention to bad decisions.
A great example of this was Mozilla's thorough and careful privacy analysis of FLoC (https://blog.mozilla.org/en/privacy-security/privacy-analysi...), and looking at Topics (https://github.com/patcg-individual-drafts/topics) Chrome seems to have spent a lot of time addressing that feedback.
Customers who match the criteria selected by advertisers can be targeted for other ads -- on different websites, streaming channels, by mail, etc. So this is how advertisers have access to information about you which they, themselves, did not collect.
Imagine the pain of selling accessories for two types of iPhones, USB-C and Lighting. The engineering that needs to accomodate two ports. The amount of people turning to the grey market to get a USB-C iPhone.
I'd be shocked if Apple ever releases two models. My bet is the first iPhone that falls under the European mandate will have no ports.
I'd give that maybe a 75% chance. I'm not totally sure they're quite ready to release a portless iPhone, which I believe would be a very unpopular change overall. Angering users hasn't always been a big concern for Apple, but I think it's more of a concern than it used to be. Wireless charging just isn't a good fit for a lot of charging scenarios.
Selfishly, I do hope that they don't release a portless iPhone anytime soon. I'd have to choose between upgrading my iPhone -- something I do every year or two -- and having CarPlay work in my car (which needless to say I very infrequently upgrade). I suppose a dongle or attachment for this purpose would be inevitable.
And the amount of people turning to the gray market to get a Lightning iPhone. I personally live in the EU, and I'll strongly consider importing my new phone if Apple decides to go with two separate models.
It's already annoying that developers assume my language based on my location, with features like this my real preference will be harder to determine.
As a user, what I actually miss is the ability to configure language preferences in the browser based on CCTLD.
And regarding the editorial comment of the submitted title:
- What it does is report only the most preferred language by default
- You can turn it off (i.e., just the language obfuscation) if you want to report all language preferences
I understand that it's easy to disable it, but it feels like a default that is strongly biased in favour of speakers of English and other widely spoken languages.
Maybe there should be a way to request all languages together and let the client pick whichever one it wants (I'm sure the text on a typical page when compressed, even times 50, is still negligible compared to the 50 MB of JavaScript frameworks it is probably pulling in) That would not sacrifice usability for privacy.
But here I see no added privacy in a normal setting if I do not use vpn, because 90% of IP addresses from Germany would report the same two languages. So only e.g. if I travel e.g. to Japan it makes me quite fingerprintable.
So I think the ideal would be if some entropy score could be displayed/predicted based on context (e.g. source or target address as above) and I could dynamically chose the trade-off between a bit of privacy and convenience.
The funny thing is that e.g. if you are in a country with a nonenglish speaking majority that has English as second language, just reporting either language assigns you to a smaller subgroup and makes no sense.
Something like
{ 0: de, 1: [en,fr], 2: [ru,es], 999: [zh,jp,...] }
where the keys are an arbitrary 'priority' score chosen by the server. 0 would be the original language, then in this example maybe 1 could be full human translation, 2 partial human translation, 999 machine translation.
The browser could keep its language preferences client-side and simply request their favourite language among the available ones.
We could replace 'breaks' with 'fixes', and we'd have the same kind of problem, but with the opposite bias: "Brave fixes language reporting in browser for more anonymity"
Related gripe: Twitter will only offer the report form (for e.g. a harmful tweet) inside Germany, in German. I do understand conversational German but not German legalese; I will not bother to select which exact subparagraph of the communications legislation the tweet violates, I will just close the tab and let someone else report it.
Union Jack? English Flag? USA Flag? Some horrible hybrid between them?
None of them is a perfect fit. As a techie I'd say ISO codes ('en', 'en-gb', 'fr', 'de') but I'm not sure how much those are understood, and probably not so hot for non-latin scripts.
This lets you specify an ordered list of whatever languages you prefer. If you haven't set it manually, usually it defaults to whatever your system language is.
In one approach, you can look like everyone else. This is hard to maintain, as any singular value can make you stand out from the crowd, making it easier to sift and isolate you. However, it is easier to implement, because you just need a bunch of constants in the software.
The other approach, is that every single time, you look like someone completely unique. Whilst difficult to get right, this approach does mean you look unique and you do stand out. But every single connection has that feature, which makes it rather difficult to get two completely unique profiles and determine if they are the same person.
https://github.com/brave/brave-browser/wiki/Fingerprinting-P...
and a bunch of others fingerprinting methods.
Even if it changed weekly, that’s valuable data. How many websites do you access in that time?
You supply the username/password, click SUBMIT button and they all return back to the same login page, but still not logged in.
- Brave Shields Up - Shields Up, all 16 combination - Shields Down
So because shield is down and login does not work, I’ve stopped all testing with Shield-related effort.
Version 1.38 (22.5.13.17)
> By default, Brave will only report your most preferred language. So, if your language preferences are “English (United States)” first, and Korean second, the browser will only report “en-US,en."1 Brave will also randomize the reported weight (i.e., “q”) within a certain range.
> If fingerprinting protections have been set to Strict, Brave will instead always report the language preference as “English,” which ensures the largest available anonymity set2. And here, too, Brave will randomize the reported weight (i.e., “q”) within a certain range.
> ...Brave users who wish to share more information about their language preferences with websites can easily configure Brave to do so. Users can disable the font / language protections by visiting brave://settings/shields and toggling off Reduce the identifiability of my language preferences.