How does one actually make themselves non-unique?
And can we make things more transparent? How many bits is enough?
How does one actually make themselves non-unique?
And can we make things more transparent? How many bits is enough?
Instead of asking "How do we become "non-unique""?
We could ask "How do we send less data to the parties who want to do tracking?"
For example,
User A sends 2-3 HTTP headers, the bare minimum, e.g., Host, Connection, maybe User-Agent or Cookie in some instances if required. User A has disabled Javascript unless she needs it.
User B sends a potentially unlimited number of HTTP headers (User B does not care to control the headers she sends, she just leaves that to the applications). User B leaves Javascript enabled regardless of whether it is needed.
Each user may be "unique", with sufficient effort from the party doing tracking, but there's an argument that User A is less interesting to the tracking parties than User B. Both users have footprints, both can be tracked, but one is sending much more data to the trackers than the other. She's leaving more detailed tracks, so to speak. User A is leaving a more generic footprint.
There is also an argument that finding "uniqueness" among a large group of "User B's" would be easier than amongst a large group of "User A's". If we were trying to achieve the impossible goal of "non-uniqueness" it would arguably be easier to try to have all users appear to be identical to User A than trying to get all users to match User B what with all the additional potential variables User B presents thanks to uncontrolled HTTP headers and Javascript's access to her computer's resources (and all the potential issues and "options" that raises).
The question is how to do it tho, because some things might actually be useful: if your browser tells the site you prefer dark themes the site can react and display your prefered color scheme. If the browser tells the site how big your viewport is it can give you a site that fills that viewport neatly — and if the site can do these useful things it can also track you using that info.
I think it’s less true if you’re trying to play a game, look over a complex dataset, do high-end image previewing, maybe host a multi-party video conference, or other highly interactive application in a web browser (for user convenience and, to some extent, because users semi-reasonably trust browsers more than random app downloads).
There is actually no technical reason that an HTTP request needs to be any more than
GET / HTTP/1.0
It's nice for the server to know your user agent, what fonts you have installed, your OS, and whatever else, but it is not necessary and so the majority of the problem of fingerprinting browsers is one created by the browser developers themselves. The original concept of the web was that the client controlled the rendering, the server shouldn't care about what specific fonts you have or the size of your screen.There is no reason that the Firefox or Safari developers couldn't decide in the next version to send only bare minimal request headers.
> nc news.ycombinator.com 80
GET / HTTP/1.0
HTTP/1.1 301 Moved Permanently
Server: nginx
Date: Sat, 21 Nov 2020 13:59:31 GMT
Content-Type: text/html
Content-Length: 178
Connection: close
Location: https://news.ycombinator.com/ > openssl s_client -connect news.ycombinator.com:443
GET / HTTP/1.0
HTTP/1.1 200 OK
Server: nginx
Date: Sat, 21 Nov 2020 17:25:28 GMT
Content-Type: text/html; charset=utf-8
Connection: close
Vary: Accept-Encoding
Cache-Control: private; max-age=0
X-Frame-Options: DENY
X-Content-Type-Options: nosniff
X-XSS-Protection: 1; mode=block
Referrer-Policy: origin
Strict-Transport-Security: max-age=31556900
Content-Security-Policy: default-src 'self'; script-src 'self' 'unsafe-inline' https://www.google.com/recaptcha/ https://www.gstatic.com/recaptcha/ https://cdnjs.cloudflare.com/; frame-src 'self' https://www.google.com/recaptcha/; style-src 'self' 'unsafe-inline'
<html lang="en" op="news"><head><meta name="referrer" content="origin"><meta name="viewport" content="width=device-width, initial-scale=1.0"><link rel="stylesheet" type="text/css" href="news.css?jacyibSTLtogl89kgrgw">
<link rel="shortcut icon" href="favicon.ico">
<link rel="alternate" type="application/rss+xml" title="RSS" href="rss">
<title>Hacker News</title>All that additional crud has probably contributed to the "relative ease" of conducting tracking as well as the "richness" of the data one can gather from tracking. Why should we ignore this simple fact.
The people behind these browers, especially Mozilla, keep assuring the public they are working to protect user privacy. This may be true to some extent but what they are not telling the public is how they are working to ensure the online advertsing industry continues to thrive, i.e., how they are working to ensure they do not upset the status quo. The words do not match the actions.
I would not rely on the browser developers to address this problem. To experiment with how well minimalism works on today's web, one can use alternative HTTP clients that allow control over headers and/or one can use a forward proxy one can remove/modify headers generated by the browser.
User A makes a single request to download bulk DNS data or Wikipedia content, then makes it available on her loopback or local network.
User B makes a series of DNS queries for specific RR's or HTTP requests for specific Wikipedia pages.
How do the behavioural fingerprints of User A and User B compare?
Trackers generally cannot see the individual DNS queries or HTTP requests for Wikipedia pages that User A makes over her loopback or to addresses on her local network. However all of User B's individual DNS queries and HTTP requests, including timings and context, are easily visible to a variety of trackers operating on the internet.
The basic principle of information theory is that these two things are exactly the same. There is no "correspondence" between them, they are one and the same thing. The amount of information that you send is, by definition, the measure of how unique it is.
I don't see how your reasoning about sending more or fewer headers fits into this; but anyhow the conclusion won't depend at all on the amount or the length of the headers, but on how many people are sending the same headers as you. It is impossible to know how much information are you sending by only looking at the bytes you send. You need to know the worldwide distribution of headers to measure that.
Your interpretation, which I think is correct in this context, seems to be with regard to the entropy of a probability distribution over internet users, and the mutual information between that and the distribution over messages. The actual length of the message is irrelevant to the math once you fix the joint probability distribution.
The argument others seem to be making is that the joint probability distribution is in fact not fixed, and that you can smear out the conditional probability over users given a message by shrinking the space of possible messages. In theory that seems possible, but I don't know enough to have any idea how well that would work in practice. If you shrank the message space to be small enough to be useful for this purpose, wouldn't that get in the way of usability?
This is not "my" interpretation, it's the standard definition of information content in computer science, as given [1] by Shannon in 1948 and used by everybody since.
[1] https://en.wikipedia.org/wiki/A_Mathematical_Theory_of_Commu...
Define the random variables M for message content and U for the identity of the user. The interpretation of "bits of information" that most people will have is H(M). The correct interpretation in this context is H(U). You seem to be confused about why people are talking about H(M) instead of H(U). But I think people correctly intuit that those aren't independent, so the mutual information I(U;M) = H(U) - H(U|M) is positive. And obviously if you change P(M), you will also change the amount of mutual information. That's why talking about sending fewer headers makes sense.
We should agree on common set of absolutely minimal necessary data (User agent, supported fonts, screen resolution and DPI if Javascript can probe those, etc) and then ALL run with the same, except on the very few sites one truly needs and which don't work with this bare minimum for some (hopefully valid!) technical reason.
We could actually package all of this into an addon or a simple recipe of settings. I'm not knowledgeable enough to pull such an effort with absolute certainty that nothing gets forgotten, but if someone does, I'd suggest "V for Vendetta" as a name, as it reminds me of the 5th of November scene.
Especially the fonts installed was very interesting and I hadn't thought of that before. The JS just iterates over thousands of known fonts in a div/span and sees if the browser can render it to build a list. It's enough that you've installed a single uncommon font and together with everything else you suddenly became unique just by that.
You're not going to block all of these "dynamic" prints just by changing browser or installing a single plugin. Even if you run in a VM, unless you actually flush and reinstall that VM every session, you're eventually going to amass customizations inside the VM that can be fingerprinted.. :/
Yet CoverYourTracks says that even my up to date Chrome User Agent is used by one in 223 browsers, which sounds hard to believe.
I haven't tested it much myself, but I suspect there's a lot to unpack here.
There's ways to do it with virtualbox and qemu as well by setting the disks to the same ephemeral style.
I wonder what already exists like this?
Even without any political interference "War is a racket"! ( https://en.wikipedia.org/wiki/War_Is_a_Racket )
1. https://www.thegreatcoursesdaily.com/synarchist-conspiracy-i...
It's highly unlikely I'm going to do anything with this thought, but if I do I think I'll call it Lignin, after the complex structural compound in plants that no organism was able to digest for the first 60 million years after its appearance.
https://en.wikipedia.org/wiki/Lignin
(I'm going to go read about mcgovern now.)
Its like if I were an analyst at a national security TLA, I’d treat TOR traffic as a big red “look at me!!!” flashing light. Sure, you _might_ be using TOR to anonymously report a pot hole to your local authorities, but you’re _way_ more likely than average to be doing something the government has “a war on”. (So as an analyst I’d assume you’re involved in drugs, terrorism, child abuse, or insisting on government accountability.)
You probably don’t want “a library of garbage”, you probably want something a little smarter that breaks all your traffic into totally plausible “normal looking” traffic but with each stream (browser tab/website pair I guess in this context) looking like a different but totally “normal” session. So your HN session looks like a totally stock Win10 Edge browser session, but when you click over to (or open a tab to) eff.org, it changes to maybe a SamsungS20 session, and when you flip somewhere like NYT - all the page loads and the hundreds of tracking pixels all see what looks like a macOS Safari session.
Do-able, I think, but more complex that simple “garbage” traffic. Needs stateful session inspection so it could do things like stripping referred headers when you change top level sites/urls, while sending them normally to image/xhr/tracking URL calls from within a site.
Various privacy trackers also block fingerprinting code.
If you need a new hobby, uBlock Matrix is the hard way to block most fingerprinting.
"All you need to do is pick up this abandoned github project, fork it, fix all the outstanding show stopper issues, bring it up to date with advances in browsers since it was last regularly maintained, and add in my own personal must face feature! Bob's your uncle! The you just need to avoid whatever inevitability it was that caused the previous maintainer to abandon it, deal with the usual crowd of self entitled and demanding-to-the-point-of-abuse users who refuse to contribute pull requests or money, and worry about tremendously overreaching to the point of fraudulent DMCA or patent lawsuits and having GitHub roll over to the RIAA or whoever doesn't like what other people use the code to do."
:sigh:
Besides that, Tor browser is really your only option.
We're in so deep it will take years to reverse fingerprinting vectors. If it's even possible
The EFF site showed two fingerprint ids that were completely unique for my browser: window size (because of my dock) and the http-accept headers (because of my language selection). That's with FFs fingerprint-resist option enabled. Chameleon can spoof those, which is great, and it gives access to the fingerprinting option which I think FF does not expose properly outside of about:config. So it should greatly reduce my identifiablity, but according to the site it does not help much, even if the specific categories are now almost unique. Like they explain, the combination seems to be the problem, or maybe they are not exposing the category that gets me.
Edit: Okay. The solution seems to be: Chameleon with most of the options enabled, so as much spoofing as possible, but without activating FF's fingerprint-resist option. Probably sending no data is worse than sending spoofed data!
+ Ublock of course.
So while Chrome is such a dominant browser? No. It’s not gonna happen. Google will keep writing “inadvertent errors” that exempt their own tracking cookies from user instructions to delete them.
Turning off javascript is a very good thing to do. The more people that do it, the stronger a protection it will be.
Random browsing => no javascript.
Play a browser game => javascript.
I have this today with Flash, would be nice to have it for all client-side code execution.
It's gotten a lot more annoying the last several years with 5mb webpages. Backing up the whitelist saves a lot of time.
For perspective, I use 4 separate profiles (--user-data-dir) listed in descending order of how annoying they are for me to use.
1). School browser: Chrome + uBlock origin
2). Shopping/low security (when the payment processor is an iFrame and I don't want to refresh...): Firefox + uBlock origin + CanvasBlocker
3). General browsing like browsing Google or YouTube: Chrome + uBlock Origin + uMatrix
4). VPN browser: Chrome + uBlock origin + uMatrix + CookieAutoDelete + VPN extension
I've gone from 2 browsers (Firefox + IE6) to 4 browsers (Firefox, Chrome1, Chrome2, Chrome3). By 2030 I'll be running 16 browsers in a virtual machine on a remote server that I connect to with my browser browser.
[0] https://www.ghacks.net/2020/09/20/umatrix-development-has-en...
[1] http://forums.mozillazine.org/viewtopic.php?f=38&t=368230 (2006)
No updates is not necessarily a bad thing. Sometimes things work well enough to leave alone.
uBlock Origin solves this in two clicks. WFM.
> Your browser fingerprint appears to be unique among the 2xx,xxx tested in the past 45 days.
> Currently, we estimate that your browser has a fingerprint that conveys at least 18.xx bits of identifying information.
Biggest offenders: USER AGENT and HTTP_ACCEPT HEADERS. Especially the USER AGENT is crazy, 9 digit browser version to everyone who asks?!
Sadly, this is of limited use. Defense against fingerprinting is like herd immunity. If everybody else already has a unique fingerprint, there is not much an individual can do to avoid being uniquely identified as well. At most one can spoof one other unique individual. Plus the EFF recommendation is 'latest Chrome on Windows' which is a moving target.
Would be nice if the EFF site in OP would recommend an agent id to spoof to, at least that would help building a small, but non trivial herd of indistinguishable users. And then a popular extension like uBlock Origin would track this agent id and set it by default for all its users.
Edit: list of top UAs:
https://techblog.willshouse.com/2012/01/03/most-common-user-...
privacy.resistFingerprinting = 1
Among many other things, it sets UA to the LTS release, and `HTTP_ACCEPT` to a vanilla en-US string.Brave does this.
I honestly could not live without it. At this point I have pretty much every news and recipe site on the internet blocked. Visiting a new site, as soon as I hear my fan spinning up I reach for the "no-JS" button and the page suddenly becomes responsive again.
I'm not quite to the "no JS as default" level but I'm close.
Google themselves admited they are an ad company rather than a search engine company. Why use a browser by a company where their main revenues are ads.
I am tempted to use Firefox more to avoid browser monoculture. But Chrome pretty much just works.
Yes. This is why it would ideally be done by the browser, not by individuals. If Safari reported only its top-level version number (and only exposed the installed-by-default fonts, and so on) then millions of others would suddenly look like me.
It's basically a virtualbox VM that you run a browser inside of.
No, it's not really practical. But yes, it works. If you're doing nefarious things, that's probably the #1 thing you should use.
Though this is highly desirable, for a guy like me who is rather paranoid I'm not going to worry too much. I disable js and block lots of sites (based on MVPS but also my own personal list - anything that gets through gets put into my personal list). After that, I don't care too much if they track me cos they aren't getting bugger all useful - if everyone did that, what market would be left for advertisers?
I'll still read more of the article and try to close more holes, but perhaps a 95% solution is sufficient? What do you think?
They promised that the tech would always give a stable and unique id from the browser. And it worked too, but it wasn't public and not for ad tech purposes.
I have a POWER9 desktop, and a fair share of other users do so for privacy/security concerns, pretty hardcore Tor browser guys and the like. I've mentioned before that fingerprinting firefox on ppc64le would be very easy because of timing the non-JIT'd JS engine. I guess there's potential for much more specific fingerprinting.
How would you defend against that profiling tech?
Brave, piehole, ubuntu, custom desktop computer. I didn't even disable javascript.
https://web.archive.org/web/20180714043311/https://iotdarwin...
crucial bit is to blend into the crowd instead of standing out (the paradox is the harder we try to mitigate on a technical level the more we'll stick out). instead use hardware compartmentalization, pseudonymous identities (not anonymous ones) and focus on operational techniques (modify your behavior). technology can be a means but often it is just part of a strategy. (e.g. instead of "let's encrypt all the comms" why not eliminating some comms altogether - no need to worry about data that doesn't get stored in the first place etc)
1. TOR browser on Safest setting
2. Firefox using the strictest privacy settings except allowing 1st party cookies, and using Private Browsing; toggling privacy.resistFingeprinting to True in about:config; and using the uBlock and NoScript add-ons
These were successful on both desktop and mobile (except iOS)
Actually, let's look at a different solution up the same vein, how does someone become NOT themselves? I think that right now it's less about how to blend in, and more about how to swap identifying features.
What about a pool of users, like a vpn, except the pool exercises identifiable information swapping!
Edit: removed typo.
And sure, N is pretty big for popular websites. But the less popular the site, the less it holds up.
Tor is about the most effective thing we have, but I'm not sure "really effective" is fair. It gets the job done sometimes, but it also leaves a lot to be desired.
Still, I'm grateful Tor exists.
Is there anything else that makes it not really effective?
I've never seen that message before, but that would explain why it always opens up 'windowed'.
[1] https://trac.torproject.org/projects/tor/ticket/31059 [2] https://blog.torproject.org/new-release-tor-browser-90
Firefox comes close with uBlock, NoScript, privacy.resistFingerprinting, and strict privacy settings