Browser Fingerprinting: A Survey
arxiv.org
arxiv.org
Stalking is unwanted and/or repeated surveillance by an individual or group towards another person.[1] Stalking behaviors are interrelated to harassment and intimidation and may include following the victim in person or monitoring them. The term stalking is used with some differing definitions in psychiatry and psychology, as well as in some legal jurisdictions as a term for a criminal offense.
This should also apply in the digital world. Doesn't it?
What about this: Noticing and remembering people's faces, clothes, tics, etc. when they walk up to you. Also you have blanketed the planet in cameras that all feed back to a giant database that builds a detailed dossier on almost everyone, based on faces, clothes, tics, etc.
The way you put it sounds almost accidental, like no effort was made, with no real goal in mind. That's not a faithful analogy of the pervasiveness of Google Analytics, Facebook buttons, etc.
It is when you place yourself in a position to be nearly every one of the hundreds of people that person interacts with in a day.
Quantity has a quality all of its own. You can't property reason about tracking and surveillance through analogies to conventional social interaction.
https://en.wikipedia.org/wiki/Cyberstalking
I don’t think you will be able to convince a court that browser fingerprinting is cyberstalking though, but you could certainly try.
The cyberstalking industry has many players. Most are small. A target-rich environment.
Shifting the investment or financing calculus, via increased investor risk, possibly including legal liability, would be huge.
Ultimately, I think it should require user consent. (And, speaking of consent, I wonder if simple fingerprinting and associating a browser with an ID would violate the GDPR. It's clearly intended to be "personal" to one particular user, but is it "identifiable" information? Does it matter?)
- User agent
- List of Plugins
- Screen resolution and color depth
- Canvas
- WebGL Renderer (~GPU Model)
Maybe Canvas can be fixed to a degree, but the others feel like things that will inherently be different. Firefox with the privacy.resistfingerprinting setting solves all of these at least partially (and more), but not without some usability tradeoffs.
GPU models also should not be given, browsers need to abstract features and performance into something else. Like if I'm writing a WebGL app, I should just be able to test a few baselines configs.
User agent is a tricky one, Chrome on Android gives too much info because it's built by an advertising company.
So no more ad blockers?
For instance, several versions of Firefox had a serious bug in shared workers, which appeared intermittently when users opened your site in multiple tabs https://stackoverflow.com/questions/51092596/feature-detecti... - I had to use the user agent to work around this.
Bundling into adblockers or browsers directly would help.
Chrome could change its header to the header of a competitor or a neural one, but has no reason to (Google's main business is delivering ads). For other big browsers like Firefox and Safari it would be a huge loss to change the header: they would vanish from browser market share statistics, and with that the business case to make websites work well with their browser dwindles. That in turn would reduce their actual market share. That only leaves the browser with a market share so tiny they are ignored by developers, and those browsers do often spoof the user agent header (or at least make it very convenient to do so).
You don't need the model to fingerprint it. It's performance is probably more than enough.
edit: also your comment seems to imply that plugins are used for websites trying to do something the browser doesnt support (ex flash) -- while this is sometimes done, most plugins add additional functionality for the user, ex adblockers, usability improvements, etc
The `navigator.plugins` used for fingerprinting is just for NPAPI/PPAPI stuff like Flash, Java, Silverlight. Just about all browsers are also in the process of deprecating support for them.
(it can also be possible to fingerprint browser extensions, but that is a bit trickier as it's not served up directly by the browser as a handy array)
The rest of these "attributes" are not HTTP headers, and occurs after the resource is fetched. We do not need to give away this information in order to retrieve the resource. The fingerprinting problem arises because the software we use to do the retrieval does far more than retrieve web pages. These other attributes cannot be detected unless the software used to retrieve the web page also includes other features that process what is retrieved, including a Javascript engine that runs code straight from the internet ("browser").
Looking to the authors of a popular browser to solve the problem they themselves have created. Is it rational to expect that this is where the solution will come from?
If one is willing to make some "useability tradeoffs" then one might use a simple http client to retrieve web pages instead of a popular browser. After the pages are retrieved, then one can use a popular browser to view them if desired; the browser need not have access to the internet. This is what I do. Not to avoid "fingerprinting" but just because it is faster and more robust.
With the help of a trivial program that transforms urls into http (I wrote one), one can retrieve the text/html of web pages, without requesting all the third party advertising cruft, with any tcp client such as netcat. HTTP pipelining becomes easy, making consumption of large quantities of information more efficient -- many pages, one connection. The two programs together form an "http client" with, relative to a popular browser, or even a program like curl, very low code complexity.
Whenever I see the "fingerprinting" debate come up, I always go through the same thought experiment: For situations where "fingerprinting" is a problem, web users all use the same simple http client. All send the same minimal headers. The web must accomodate users not the other way around.
Maybe if you have script disabled, but if you have javascript it's trivial to detect which browser you're on based on javascript implementation quirks alone.
>We do not need to give away this information in order to retrieve the resource.
You clearly haven't seen mandatory javascript sites.
You are making the assumption that I am using a browser to retrieve the page. I can use a simple http client with no features such as JS to retrieve the page. I have full control over what headers I send and the contents of those headers.
"You clearly haven't seen mandatory javascript sites."
I have seen many sites that prod the user to enable JS yet still return the same text/html even when JS is disabled. Sometimes I am only after the content of the page: text, URLs pointing to images or other resources, URLs pointing to other sites, etc.
I have seen some sites where there is "no content" in the initial page fetched, with or without JS enabled. These are dummy sites, hollow shells, skeleton pages. The technique is simply misdirection. All the useful content is pulled from another site via a second request (JS triggering the request). Maybe this is a CDN. I simply read the JS or use a debugger to find this "endpoint". Then use the simple http client to request the content as usual. These dummy sites are not the majority of sites on the web, nor even the majority of sites posted to HN.
The endpoints supporting these "no content", dummy sites often deliver the text with fewer or no tags, without HTML. Maybe it is JSON. This makes it easier for me to format it the way I like it. What many HN readers may perceive as a growing nuisance I believe could be a blessing in disguise. I prefer delivery of content in a more "raw" format allowing the user to process it as she pleases client-side, with the software of her choice. This leaves no room for advertising. It is more efficient and the user becomes the creative one, deciding how she wants to present the content to her own eyes.
If so, I want to shake your hand, sir.
For non-commercial use of the web, I use what I want to use. A simple http client and/or text-only browser.
My day-to-day use of the internet is almost always 100% non-commercial.
I never want to use a Javascript-enabled browser for simply retrieving information from the web. For example, reading. This is non-commercial activity.
For simple information retrieval, I find a large, complex, popular, recommended "browser", a single, "do-everything" program, is overkill and, counterintuitively perhaps, such browsers loaded with "features" actually limit what I can do. Using smaller, simpler "one-purpose" programs to retrieve web pages allows me more flexibiilty. I can be more productive.
Using single-purpose programs, I also find the retrieval process to be more reliable and robust, not to mention more transparent. I know exactly what I am sending. Unlike the popular browsers, these programs are not fetching and running code from the internet automatically. I feel a greater sense of control.
Canvases should never be readable by the other end without explicit permission. Basically transmitting anything rendered to a canvas should be equivalent to using my web camera from an integrity standpoint.
Time zone, operating system (from a short list) and language is fine, because it’s not a lot of entropy and my ip is known anyway unless on a vpn.
Gl renderer/graphics card, desktop resolution - nah. The viewport dimension is all that should be known.
Basically the browser should by default have an entropy budget for my entire fingerprint that ensures I’m not personally identifiable.
This could be fixed with a bit of courage. Firefox could use "Firefox", Safari "Safari," etc. UA sniffing and UA strings are completely absurd at this point, and need to go away.
> - List of Plugins
JavaScript should not see this, especially not their versions.
The last three are harder to avoid for "responsive" web design. One attribute you didn't mention was the font list. There is no good reason for JavaScript to be able to see this. The model should be that the page requests a font, and the browser uses either that or its best try at a substitute.
If Firefox tried that, it would instantly become that weird browser that makes Google look bad. The internet would be flooded with tips on how to set a custom UA that made it pretty again.
I just tried, and it's great! I had forgotten that search could be so snappy and simple. The more likely outcome, unfortunately, is that Google would adjust its UA-sniffing to serve the same 2019 bloat to "Firefox."
You can just measure performance of those things then.
Chrome on Android reports your phone model and build number. If you're using a lessor known phone with a carrier specific ROM, then you're in a really tiny population.
WebRTC is another one which you can disable. JavaScript, in general, as well, requiring uMatrix.
Chrome leaks the extensions you use, Firefox does not (AFAIK). As for blocking canvas, see [1]
To test your fingerprinting I can recommend IPleak [2]
The problem of this all is, its a usability nightmare.
[1] https://addons.mozilla.org/en-US/android/addon/canvasblocker...
Sure... "leak", wink wink.
If you use a common device and browser kept up to date to the latest versions, then you are not going to be unique. Using Safari on an iPhone will make you similar to 30% of web traffic.
Browsers also update all the time changing the user agent so it isn't that reliable tracking over extended periods of time.
The am I unique websites seem to compare your user agent to a database with outdated browser versions from the past data, making your score look much scarier than it really is. If you keep things up to date, chances are that you are not that unique.
Basically Canvas is a JavaScript API for programmatically drawing graphics and text. Subtle differences in color, font rendering, antialiasing etc will produce different rendering on different hardware depending on OS, browser, GPU, drivers, etc. This lets a site running JavaScript to generate a hash unique to everyone with your specific combination of software and hardware and uniquely identify you without cookies.
EDIT: typo
However first you need some serious regulations to split Google and take away its browser "business". As it is Google is doing everything to preserve and advance tracking and fingerprinting in web browsers.
Other than that, things like presence/absence of common variants like bold, and possibly some common fonts come packaged with office software rather than the OS.
The user agent for mobile safari doesn't identify the iphone model, only that it's an iphone[1]. Knowing the precise model definitely helps to fingerprint more.
[1] random search: Mozilla/5.0 (iPhone; CPU iPhone OS 12_1 like Mac OS X) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/12.0 Mobile/15E148 Safari/604.1
Especially as it isn't as simple as a standardized list of fonts when the Canvas hash is 100% unique for everyone.
Still very disappointing.
The issue will be solved when all tracking companies have collapsed and that industry is dead. Clearly we aren’t there yet.
Here's a quote: "There will also be new security measures to prevent digital fingerprinting, or the use of things like installed fonts and plug-ins to help track users across the internet even with privacy settings active. Websites will be given a stripped down, simplified system configuration so every user's Mac looks like every other user's Mac."
Andale Mono, Arial, Arial Black, Arial Hebrew, Arial Narrow, Arial Rounded MT Bold, Arial Unicode MS, Comic Sans MS, Courier, Courier New, Geneva, Georgia, Helvetica, Helvetica Neue, Impact, LUCIDA GRANDE, Microsoft Sans Serif, Monaco, Palatino, Tahoma, Times, Times New Roman, Trebuchet MS, Verdana, Wingdings, Wingdings 2, Wingdings 3 (via javascript)
I have hundreds more fonts installed. Panopticlick says that about 1 in 14 browsers have this value, so your browser should have it too.
So that part seems to work, I was just confused because the effect on overall uniqueness is very low. Still I applaud the efforts by the WebKit Team.
Depending on the strength of defence required, anything from a low-cost registration fee (see Metafilter) to some form of recommendation-and-vetting or simultaneous in-person token-granting event (simultaneous to avoid entities being in two places at once). And auditing for abuse.
There have been several proposals for such systems also assuring some level of anonymity.
It's unlikely any approach will be perfect, though an arbitrary level of assurance is likely possible at some cost.
Note that surreptitiously fingerprinting and preemptively certifying to establish entity uniqueness are vastly different from a user awareness and concent perspective.
Sure, someone can create another account using a new browser, within a VM, from another computer, inside a VPN. It's all about making it much harder. If the primary use of fingerprinting was to protect community from bad actors, like those violating a set of community guidelines, then maybe the extra effort it would take to get around those issues might give them enough time to diffuse.
You could use IP address, although that only works if the user isn't on a public / shared network. It's also easily bypassed by spinning up a VM on a cloud service provider and using an SSH tunnel.
Since you used polls as an example: StrawPoll.me [0] is an online poll site which lets you select different duplication checks based on your requirements. The choices are: IP, browser cookie, none, or require user sign in. They also give you an option to add a CAPTCHA.
* Cookies: Can be deleted * IP-Adress: Not unique, because ISPs rotate them; also VPN * Login: Well create a second one * Methods from Universities using nth letter of name and nth digit of birthdate: Just make up a new name.
Sorry - but unless you are using an analog medium or asking the questions in person the numbers can be inflated and there is no way to have 100% data quality.
But in most contexts this is ok. So I would probably go the most easy way: Cookie.
I would - at least not in the European Union go with fingerprinting and such stuff, as I am not sure how this plays out regarding GDPR as this would be PII you are storing.
Nevertheless, what needs to happen is that all major browser makers come together and simply create a set of standard API values that do not harm daily browsing and make it possible for users to blend in with the masses, if they opt-in to activate
It would be sufficient to create a couple of uniform user agents, list of fonts, list of plugins, canvas hash, platform and webgl data to bring the uniqueness down.
Also, living in the EU won't protect you from getting fingerprinted.
True, but most browsers have an auto-update scheme, so basically everybody is on the same version at a given time (by approximation).
Primarily for companies that develop shadow profiles of users but also companies like google and cloudflare where their tracking is a result of an opt-in by the site operator.
Would it be difficult to prove I am in fear (rational) as a result of their stalking?
I'm not aware of the law in this area. I'm a lawyer, and I could at least see a plausible argument, but the only way to know for sure would be to try to sue them, or find other cases where someone has successfully done so.
I'd like to see an organization that checks for this, and gives fines, and perhaps even withdraws the right to use a domain name. And I'd like to see more responsibility with site owners for using third party code.
By the way, I think a withdrawal of the right to use a name (brand) is a very appropriate way to penalize a serious privacy violations. Brands are all about customer trust, and if that trust is violated, then it seems to me only fair that the right to use a name is taken away.
Prevalence:
- 1.4% for canvasfingerprinting
- 0.325% for canvasfont probing
- 0.0715% for WebRTC
- 0.0067% forAudioContext
The above were found only on the most shady of all websites, and good content blockers block all those scripts.
My bet is GDPR was the death blow for this kind of scripts. For small companies without a room full of lawyers data has become a liability.
So fingerprinting is now basically in the hands of google, amazon, etc.
If it was useful, it would be much more prevalent but too many people have the same fingerprint.
IP address is a better method of tracking for comparison.
The industry used fingerprinters, but for one it didn't really help them make more money (because you want to track users, not systems), and there was a big backlash.
According to whom? My browser is unique according to panopticlick.
Curious, was the lawsuit in that study also based on GDPR?
So far I haven't seen much coverage of GDPR based lawsuits in the media. Also, I'd like to know if there have been studies that show that EU residents can now (after GDPR) browse the web without leaving a trail of information.
Quote about canvas fingerprinting:
"Comparing our results with a 2014 study [1], we find three important trends. First, the most prominent trackers have by-and-large stopped using it, suggesting that the public backlash following that study was effective. Second, the overall number of domains employing it has increased considerably, indicating that knowledge of the technique has spread and that more obscure trackers are less concerned about public perception. As the technique evolves, the images used have increased in variety and complexity, as we detail in Figure 12 in the Appendix. Third, the use has shifted from behavioral tracking to fraud detection, in line with thead industry’s self-regulatory norm regarding acceptable uses of fingerprinting."
[1] G. Acar,C. Eubank, S. Englehardt, M. Juarez, A. Narayanan,and C. Diaz. The web never forgets: Persistent trackingmechanisms in the wild. InProceedings of CCS, 2014.
Other references:
[10] W. Davis. KISSmetrics Finalizes Supercookies Settlement.http://www.mediapost.com/publications/article/191409/kissmet.... [Online; accessed 12-May-2014].
[15] Federal Trade Commission. Google will pay $22.5 millionto settle FTC charges it misrepresented privacy assurancesto users of Apple’s Safari internet browser. https://www.ftc.gov/news-events/press-releases/2012/08/googl..., 2012