Anonymous Browser Fingerprinting in Practice
valve.github.io
valve.github.io
I wrote several implementations of passive device fingerprinting for fraud/security firms and they are implemented across a broad range of industries and companies. I know quite a bit about the shortcomings of these and the challenges faced, in case someone is interested.
Also, just FYI, the big use these days is actually NOT in security or fraud, though most of the companies selling these still say that is the case. The use is in advertising. Most of the companies selling these still call themselves "security" or "fraud" companies, but the revenues from advertising are far outpacing the former categories.
For instance, security doesn't get much attention and people view it as a necessary evil. Fraud is about loss prevention in most cases, so the way you show your value is how much you save the company in loses. Your revenue is typically a fraction of those savings.
Most companies using tracking in advertising are still operating at a loss or have fuzzier ideas of "revenue". TBH, advertising is weird. People spend money like water without really knowing what they are getting in return. This obviously won't last, but there is money to be made in the meantime. At least that is my experience.
They probably use mostly public / school computers to access the website.
For example - if you are on espn.com then espn.com and all the trackers referenced from their pages would see a set of cookies and variables specific to espn.com. Browse over to icanhascheezeburger.com and even if the lolcats use the same trackers as espn.com does they would see a different set of cookies and a slightly different set of values - like browserid incremented by a .1 or a couple of non-essential plugins missing from the list of installed plugins.
I welcome any technical feedback on the idea.
"We ran our algorithm over the set of users whose cookies indicated that they were returning to the site 1-2 hours or more after their first visit, and who now had a different fingerprint. Excluding users whose fingerprints changed because they disabled javascript (a common case in response to visiting panopticlick.eff.org, but perhaps not so common in the real world), our heuristic made a correct guess in 65% of cases, an incorrect guess in 0.56% of cases, and no guess in 35% of cases. 99.1% of guesses were correct, while the false positive rate was 0.86%. Our algorithm was clearly very crude, and no doubt could be significantly improved with effort." (any typos here are likely the result of copy/pasting from the pdf)
The easiest way, tbh, is to look at patterns for that fingerprint(s) rather than assume they are all unique. This is where a rules engine comes into play.
The "real" way to do this is to track upgrading components (like plugins etc) and create new fingerprints as they evolve, associate them with the old ones and over time sunset old fingerprints. This is harder to do in practice, though.
Also, some fingerprinting approaches track things like monitor resolution, which is tough b/c a docked laptop has a different fingerprint than the same undocked laptop. So the approach differs based on what you track.
Of course, if the plugin is well known, it could be removed by the fingerprinting code, unless it can disguise itself as something else -more or less random- when it is installed.
Authentication systems using Browser ID create a value based on the browser identification bits and apply a ratio of how close the value is to a given user's fingerprint, so your values may change (like when you resize your screen) but it'll approximate the fingerprint as the user's. In other words, if it's close, it still matches.
As for mobile devices, here too I always install my preferred software, so I don't see how mobile devices have lower configurability. That said, there exists databases with which one may at least identify the mobile device in order to serve up content appropriate for that device.
Alas, my tendencies make my devices uniquely identifiable!
(From Panopticlick) Your browser fingerprint appears to be unique among the 3,140,429 tested so far.
Currently, we estimate that your browser has a fingerprint that conveys at least 21.58 bits of identifying information.
That's for FreeBSD running Xombrero.
That being said, people who tend to be worried about security also tend to have more unique computing configurations and thus would be more identifiable. Interesting.
Most implementations, though, make the assumption that if they get JS info, that info is good. Better to give them wrong/inaccurate/changing JS info than no info.
So, I leave JS on, use some personal plugins to change variables and give them bad data on each request. They won't know and will just get bad fingerprints that are not seen again.
See this comment: https://news.ycombinator.com/item?id=6046810