Web fingerprinting is worse than I thought (2023)
bitestring.com
bitestring.com
A couple of years ago, Apple launched App Tracking Transparency as a way to reduce tracking across their iOS app ecosystem. People predicted that this would be devastating for companies like Meta and Snap, and it was -- briefly, for Meta. But Meta seems to have rebounded very quickly, maybe Snap not so quickly. The rumor I've heard is that Meta threw every brain they had against the problem of finding new ways to track app users, which presumably involves some similar type of fingerprinting. The revenue success strongly indicates were successful. But if this is true, nobody has much written about it.
They found sneaky ways on Android. There is no way they aren't trying to do so on iOS. One must always assume malice with anything Meta.
Perhaps that's not "bad enough" but I think the general sentiment that corporations value profits over harms to people (especially since they often try to distance themselves by offshoring, etc) applies to Apple as well.
Meta and Google make their money primarily from advertisers, Apple makes money from consumers buying iPhones. One of the upsides to paying for something is that the company is incentivized to keep you paying or get you to pay more.
Something I remind people who buy cheaper Android phones and then complain about ads - the OS development is being subsidized by those ads. From Google’s perspective, securing their revenue stream is the justification for Chrome and Android’s existence. It’s not a purely altruistic move to fund their open source development.
Charts of the revenue stream for some major tech companies:
https://www.visualcapitalist.com/charted-how-does-meta-make-...
https://www.visualcapitalist.com/alphabets-revenue-breakdown...
https://www.visualcapitalist.com/charted-how-apple-makes-its...
https://www.visualcapitalist.com/how-amazon-makes-its-billio...
https://www.visualcapitalist.com/how-microsoft-makes-its-bil...
Older aggregate chart:
https://www.visualcapitalist.com/how-big-tech-makes-their-bi...
Also, WhatsApp refuses to be usable without giving it Contacts access. I had to use the app, login to the web client, and then I was finally able to type a phone number to start a new chat.
I ended up uninstalling it, but there's plenty of people AND business that nowadays mainly or even only use WhatsApp that it's painful to be on the privacy-first side.
https://developers.google.com/identity/sms-retriever/overvie...
Here's an example link,
For WhatsApp, WhatsApp business lets you easily start conversations just by entering any phone number. But yeah it’s still WhatsApp and meta, I personally avoid it as much as I can.
I don't fault you for not trusting Meta - I feel the same.
That said, what you're talking about here is an OS feature nowadays.
For example,
1. Export contacts from the Contact app to a file if it is not a new phone
2. Disable Contacts app
3. Install a different contact database such as OpenContacts from F-Droid or Github
4. Import contacts from the file into OpenContacts
WhatsApp will not import the contacts in the OpenContacts database
Further, no other app will import these contacts either
This solves the "access to contacts" issue
MACs are always randomized, even when connecting to the same network. At least as far as modern devices go.
Am I wrong?
I've had the same prefix for five years now.
And yeah, sure, my device cycles through ephemeral IPv6 addresses often, but always within the same prefix.
Group IPs somewhere between /64s and /56 and you'll essentially get a household identity, at least for a few days to a few years.
Proper randomization can be enabled through the developer settings.
https://source.android.com/docs/core/connect/wifi-mac-random...
MAC addresses don't leave the local network, so it's not relevant to web tracking. Moreover it's randomized by default on ios/android so the tracking potential is limited.
A strong relationship to Apple and cross-value marketing.
Surely these rules only apply to middle sized and smaller companies. We've seen Apple get caught bending the rules for big players, even if they don't admit it.
My "rugged" browser for regular browsing has plug-ins that randomize all this data.
Browser Plugs Fingerprint Privacy Randomizer
Clear URLs
[I don't care about cockies]
Privacy Badger
Random User-Agent Switcher
Temporary Containers
uBlock Origin
Canvas Blocker
NoScript
Font Fingerprint Defender
Not all sites will work with it. For banking and plan ticket booking, I always recommend a separate, but major (e.g. Chrome) browser without any plug-ins.
Don't bother. User agent spoofing is easily detectable and it's trivial to figure out your real user-agent based on js implementation differences or TLS fingerprinting. All this does is get you banned/flagged by security vendors, on top of sticking out like a sore thumb.
>Canvas Blocker
>Font Fingerprint Defender
Also easy to easy to detect because randomized values will put you in the bucket of "uses privacy extension" users, which is probably a smaller bucket than whatever hardware profile you're on (eg. macbook pro m3 14").
>>Random User-Agent Switcher >Don't bother. User agent spoofing is easily detectable and it's trivial to figure out your real user-agent based on js implementation differences or TLS fingerprinting.
JS is blocked by default on my browser.
>Canvas Blocker >Font Fingerprint Defender
> Also easy to easy to detect because randomized values will put you in the bucket of "uses privacy extension"
Hm. How are they going to detect it is randomized? They would have to identify me first again as the same user and then conclude I randomize these values.
The major browsers can still be differentiated via default headers and TLS fingerprints, none of which requires js. Moreover if they're inconsistent you'd get flagged with "spoofs user agent", which makes you more identifiable than something like "firefox on mac".
>Hm. How are they going to detect it is randomized? They would have to identify me first again as the same user and then conclude I randomize these values.
Because a given canvas/font metrics value should return the same result given the same graphics hardware/font set. If you randomize the results it basically guarantees that your fingerprint has never been seen before. This might seem like a good thing (because you're randomized every time), but any competent fingerprinting implementation is just going to flag you as "spoofs canvas/font information". The point isn't necessarily to identify you as any particular user, it's to use the fact you're spoofing canvas/font/user-agent to fingerprint you further.
Also, its unusual enough that its unlikely they will bother trying.
All of this is overkill anyway unless you actually think you’re up against a determined actor targeting you personally. If you are, they will bother trying.
How do they know they are randomised rather than actual properties?
(Or it can just use properties of the extension like monkey-patched function toString() outputs to identify its users, which, again, narrows it down to a very small group.)
Yes! You are unique among the 4162649 fingerprints in our entire dataset.
Two visits...
Yes, fingerprint.com realizes that I am the same visitor. But ONLY IF I access it from the same IP address. This is impressive, but in the end not so much. They claim VPN does not matter for them. It does. Probably one of the last things that makes my browser identifiable.
QED...
On stock Mac OS Safari (no plugins, no hardened config), I did what they asked and visited their site in incognito mode via a VPN. It gave me a different id, with a message gleefully announcing that "your ID is the same when you're in incognito mode!" It even showed me some supposed visit from a minute ago.
Jesus what a scam.
Money always finds a way. Everyone thought the changes made a few years ago would hurt Meta but they make $70 billion net profit. At a minimum, they only need a good relationship with advertisers, and a (sort of measurable) increase from a campaign. Also ads are different now. One address may see the same five seconds of an ad hundreds of times. That is a much easier ecosystem to correlate targets through data enrichment.
Meta hoovers up every detail because they can. Knowing that user #7227724 spends 23 minutes a day in Spotify might make the ad targeting 0.4% more accurate, but does not seem like the lynchpin for the entire business.
The old site had a blog post [0] where they explicitly said they were using fingerprinting, and even called it "privacy-compliant".
I'm sure they're not unique in the service they provide, but that was the first time I'd seen someone brag about browser fingerprinting.
[0] https://web.archive.org/web/20240527125312/https://www.reven...
It's pretty hilarious legalese and tells you nothing about what it even achieves. Maybe makes you a Very Important Marketing Target.
One thing that struck me was the 'Under penalty of perjury, I declare all the above information to be true and accurate'. Shame they seem to require validating request by email. It'd be fun to take a PII breach and throw all the emails you find at 'em.
Firefox, VPN, UBlock Origin, Privacy Badger, and UMatrix plugin to block cookies and javascript by default. (You can easily whitelist first and/or third-party cookies and/or JS on sites of your choice.)
Actually, usually librewolf instead of firefox, but not a big difference I suspect.
The unsolvable problem is that modern websites are not simply documents but rather full-blown software with web browsers their runtime environments, and you simply cannot enable that amount of power without also enabling the power to fingerprint that runtime environment and thus fingerprint the user.
In fact, doing so will often times end up bringing donations from relevant industries directly to your opponent.
Now, this system of perverse incentive and legal bribery should be fixed at the constitutional level but thats a gigantic can of worms.
In the current system there are two methods that can circumvent the issue. The first is one deployed by the likes of Elizabeth Warren; run your campaign on a broad array of "fighting for your constituents" and don't get specific until you see already elected and drafting a bill.
The second path is underutilized and should be done more: lie out your ass to the moneyed interests. Take their money, make them promises, eat at their fancy dinners, befriend them, laugh at their awful jokes. Then just fucking dunk on them in the legislature, as quietly as possible. Make a big show of being forced to, keep the charade going as long as possible.
The inverse of this has been done a lot recently, with Sinema, with Fetterman. But the good version is quite rare, and a good opportunity to make our country a better place.
Key notes: tough to do in bigger positions because they're rarely the first public office seats people hold, so track records build. Tough to do in many districts because voters can be rubes who actively agree with the corporations stomping on their nards. Tough to do if you make too large of a profile(not really a concern).
Proof of Domain Expertise: Name the famous presidential campaign which focused directly on combating "this system of perverse incentive and legal bribery" as its core campaign message.
Edit: Hint: primary, lots of votes, lots of money
MD5 of answer: 1c02462874398d776ff28aeed2d056b1
Unfortunately this is a challenge with regulation; companies find a way to break the spirit of it as much as possible while following the letter. It's better that companies need consent to track us than not, but consent managers are dark patterns designed to deeply annoy us at the prospect of saying no.
I think if it's all client-side, not logged or retained, and is not transmitted to third parties, it should be fine.
IANAL
The GDPR isn’t the complex legislation monster people make it out to be, but for the most part common sense about handling sensitive data.
If it stores it and uses it for matters different than what explicitely advertised when you consented to it, than yes it's even worse.
edit: just saw that's a service they resell. So yeah it is against GDPR
I suppose one technical mitigation might be a permissions dialog when a script requests access to a high-risk API like canvas or WebGL. But that's unfortunately something that won't work for most users, who will just click through the dialog.
Not to mention the big players on the users’ team in the technical arms race (google, ms, apple) are also advertising companies.
By all means let’s solve it from the technical side - but also lets regulate privacy so everyone gets it not just people paranoid/technical enough to use the latest/best privacy respecting tools.
“If done right” is doing a lot of work in that sentence.
The way hypothetical regulation is spoken of in abstract terms where it’s perfect, solves everything, and everyone complies perfectly is at odds with how regulation works in the real world.
They try to balance keeping corporate donors happy with keeping people happy, and create regulations that are toothless empty gestures that only serve as employment opportunities for lawyers and consultants.
So yes, “if done right” is doing a lot of work. But i refuse to cede gov to the corps and retreat to anarcho-capitalist ideas like “this is a technical problem”. We attack on all fronts - regulation and technological solutions.
To whatever degree this is, indeed, a technical problem. There’s a simple choke point that is being intentionally unutilized.
It isn't trivial to craft legislation to separate these use cases, but it also is far from impossible if there would be political will to do it.
I think the latter is far more interested in surveillance of users where tracking is one building block.
And of course legislation is needed to criminalize tracking without user consent. It would just be an internet stalking law being applied.
Plus, fingerprinting tech would get developed for criminal organizations or intelligence agencies anyway.
whether this is justified is of course subjective
> Chromium (Chrome) is built by Google, an advertisement company which tracks its users for showing relevant ads. So naturally it doesn’t have any inbuilt protection against fingerprinting.
You could compare it to the concept of security by obscurity which is obviously bad.
Two questions jump to mind:
Why isn't this the default in Firefox?
What is the downside? I.e., what can break by enabling this parameter?
Here's what the settings do and what sort of side-effects you might experience:
> This setting may cause some websites to not display content or work correctly. If a site seems broken, you may want to turn off tracking protection for that site to load all content.
Some sites use light fingerprinting to provide features
Just of the top of my head:
- Timezone is set to UTC which means any web calendar input becomes confusing at best
- Canvases turn into random stripes, which leaves artefacts all over many websites
- Some websites outright block you as bots (twitch does this)
- Some web APIs break, which can be a pain if you're web apps that rely on them
You can add websites to a whitelist to avoid the downsides on some sites (privacy.resistFingerprinting.exemptedDomains) but it's a pain to do that for every website.
and the worst part is that this didn't changed the fingerprint generated by mentioned here site just increases suspect level to 9
After a while you develop and intuition for which browser to use depending on what you will be doing.
???
It definitely does. Are you talking about how it doesn't change between subsequent visits?
resistFingerprinting does seem to work against fingerprint.com in my experiments after clearing its website data and a browser restart.
I've had more issues personally with resist fingerprinting making major sites completely unusable (drupal.com, walmart.com..)
> For example, websites can see web browser version, number of CPUs on your device, screen size, number of touchpoints, video/audio codecs, operating system and many other details
If, for example, I upgrade my web browser in two weeks (i.e. I get a new version number), doesn't that mean that the site has lost me?
Sites like https://coveryourtracks.eff.org seem to focus on how unique your fingerprint is, but doesn't it also matter how stable it is over time?
To get a feeling for this, try: https://abrahamjuliot.github.io/creepjs/ ; https://bot.incolumitas.com/ and https://amiunique.org/fingerprint
Combined with super cookies (https://blog.mozilla.org/en/internet-culture/mozilla-explain...), that's a lot of data points to stitch together a high confidence fingerprint.
Although not perfect, FF is much better out of the box at limiting the leaks than chrome.
- Safari
- Safari private mode
- Chrome private mode
and it was not able to identify me across those.I then tried
- Chrome (normal, non-private mode)
and it did identify that as a repeat Chrome visit.Does Safari have better privacy than Chrome?
The user base won’t even be there anymore.
This is not just because people will be retiring old Intel systems, it’s also because Apple’s marketshare exploded when the M1 chip came out, so a very large portion of the userbase never owned an Intel Mac.
BTW, the article is incorrect that Chrome doesn't allow for user agent modification or other fingerprint resistance; you can: https://developer.chrome.com/docs/devtools/device-mode/overr... and there are extensions for more convenience. The article is also incorrect about third party cookie leakage from ads but it was possible to sniff the session ID in some cases, back a decade ago before everything went cookieless and dropped session identifiers from the protocol entirely. However, it is possible for advertisers to parameterize their campaigns and analytics to such a detail that they can link demographics to their internal user IDs, though it's against policy it is easy to go unnoticed. And things like location exfiltration in too many Android apps, I'm not trying to give Google a complete pass on privacy but it's clear the author made some assumptions based on bias.
Back to your question, though, there are other things you can use as part of the fingerprint. The fonts that are installed are a proxy for which applications have been installed. The artifacts at the edge of text rendered onto a canvas can indicate which graphics chip and drivers are installed, sometimes with differences even within the same GPU model and driver version. Touch tracking can tell whether you swipe with your left hand or your right hand. Timing signals can indicate CPU specs and even hint at whether you're in a VM or behind a VPN, etc. There are more, accessible from JS in most cases, and really most of it is more reliable than what's in the user agent string.
Maintained by a Brave employee, though the site is fully open in all senses of the word, as far as I'm aware.
This is on Android, so Brave is using their own browser engine, so I don't think things will be different on desktop.
I've pushed back any attempts for any kind of tracking for business purposes (e.g. fancy charts).
* ja3 seems to be slightly better, ja4 sometimes groups too many "people".
Edit* Title also needs (2023).
Edit: Have also set all other `fingerprinting` bools to False. uBlock, uMatrix, Privacy Badger installed.
Under Settings → Privacy and security: Enhanced Tracking Protection = strict. Tell websites not to sell or share my data. Delete cookies and site data when Firefox is closed. Enable HTTPS-Only Mode in all windows. Enable DNS over HTTPS using: Max Protection – I still can be detected.
Edit 2: Just tried with Brave, strictest settings. No effect, I am detected.
Edit 3: Tor works.
All machines would have 16 cores and 32GB ram, running windows 10, and 1 point-touch or mouse. And the resolution would also be fixed as reporting, and only on client would change.
The user-agent should be acting on our behalf. So, why isn't it (Firefox, TBB) utterly lying and acting in our interest? We know why Chrome wouldn't.
Tor also gave up this web fingerprinting fight without even really trying. Editing the JavaScript calls to consistently lie the same way was "too hard". https://m.youtube.com/watch?v=3wlNemFwbwE
Infantile developer behaviors like disabling paste in the password field? Or bona fide on page text that cannot be selected in the browser window?
There is no reason for Firefox to enable or honor these requests.
I've seen people argue with a straight face that these tools and their reports don't run afoul of GDPR/CCPA because they don't involve information that a user gave you on purpose, so it's not protected. Ghouls, all of them.
Not a great move imo
This doesn't look to be among the available toggles, and I hope that changes. I realize the light/dark setting is a data point for fingerprinting, but it's also something I have a genuine strong preference about.
The brave shields setting section also has an option for blocking scripts, which may work. It prevents the demo from being able to show an identifier for the user at all, but I'm not sure if it's preventing identification or just preventing the displaying of the identification.
Guess which company is coincidentally is the world's largest advertiser, largest ad broker, largest data tracker and owns world's most popular browser?
Disabling them globally means a broken browsing experience.
Conceivably you could develop some sort of heuristic that detects when a script is simultaneously poking at a whole bunch of APIs associated with common fingerprinting techniques (canvas capabilities, WebGL, screen size, installed fonts, etc) and then kill it. But it is certainly much harder than blocking cookies.
It happens client-side. Browser headers sent through for requests aren't enough for fingerprinting.
on edit: better clarify, I mean if you are fingerprinting, but not storing in such a way that you can actually identify someone (although not sure why you would use fingerprinting then) then I don't think there is a case.
programmers really have a hard time understanding the law, how does any violation of the law ever get found out, or any law enforced? Generally someone says hey this company is doing X, and then the government gets a warrant to say let us look through your stuff to find out if you are doing X.
As a normal rule most companies work something like:
"excuse me, we have reports you are doing X"
"Not exactly, this is what we are doing - we call it X1, which is why we are totally ok under the rules governing X. Our legal dept. can totally explain"
Court case instantiates.
If the company is doing something that they will actually say "no we are doing nothing of the sort!" then it is likely someone in the company will at some time say "hey they are really doing X" and then the warrant thing I discussed first happens.
At any rate finding out enforcing things can happen without perfect technical access to everything, that's how justice systems have managed to work for centuries.
Cookie banners were never about cookies or privacy. The industry designed them with some very explicit goals in mind: to force users to opt-in to pervasive tracking, and to blame "how unusable web has become" on GDPR
I was wondering why can't browsers just fake the hardware (assuming that is what it is using to recognize)? I understand sometimes these javascripts run some type of algorithm to detect how fast it was processed to fingerprint, but even those could potentially be faked by the browser. Is anyone working on such stuff?