Tracking users via CSS
underjord.io
underjord.io
The only things I can think of are how a user interacts with a page - I don't particularly think this is too concerning - although as with all these things there are possibly much more creative uses of it that I haven't considered.
There's a new image property loading="lazy" which generally will load an image when it approaches the viewport. This could also be "abused" in similar ways.
If this does turn out to be a privacy concern, browser settings/privacy addons could simply load all lazy images or images refered to in CSS/JS files on load which would nullify this technique.
I don't think you can do anything particularly nasty even with CSS variable programming which can apparently be used for interactive games (https://github.com/propjockey/css-sweeper). As I looked into things writing for this post I couldn't come up with a non-JS way to transition much significant data into CSS.
It would seem to me that the only advantage of this technique is to track mouse movements on the page in a very low resolution, and likely labour intensive way. Low resolution because it only works the first time on each page load, and is going to be tricky to get any meaningful granularity on the data.
I am not particularly concerned about this but my privacy concerns are definitely lower than the general consensus of this site. I only compare to HTTP logging as this is the most hidden/covert way of tracking users.
* Track what fonts a user has installed by asking for preinstalled fonts and using loaded fonts as a fallback
* Track screen width and height by conditionally loading an image
* Track screen height vs window height (linked to OS and various user settings)
https://github.com/jbtronics/CrookedStyleSheets https://news.ycombinator.com/item?id=16157773
This is less interesting than existing JavaScript techniques to identify or filter out crawlers. It works, when the image is not cached, but it's fundamentally inferior to anything you can do with JavaScript, and if you don't have JavaScript at all, I don't see why you would want to care. Just lump that little bit of traffic in with the bots for analytics purposes.
The big advantage of this is that you can still track and uniquely identify people even if they turn off cookies and javascript.
Except that many people don't want others to know how much they read something... I don't care, but many believe that is private.
> If this does turn out to be a privacy concern, browser settings/privacy addons could simply load all lazy images or images refered to in CSS/JS files on load which would nullify this technique.
So you load more invisible pixel and make this even more effective as now you'll be able to get much more granular data, like the scroll position!
Much like Nigeriam Scams self-select the most naive users with their silly stories, some advertisers may likely get better impression/click ratios once the savvy users are out of their game.
Do the tracking inside the browser, send it back, render the ad server side and send it to the user?
It's already being implemented in many places.
That said, I think minimally-tracked (i.e based on what the website you're using knows about you) server side ads would be a great step forward for both user privacy + ethical content consumption, I just don't see advertisers jumping at the opportunity.
User goes to news.com, which loads ads from adserver.com. Adserver.com sees the cookie, and shows the ads about shoes.
If you do everything on the server side, shoes.com won't be able to place a global tracking cookie. So when the user goes to news.com, they'll have to see the generic ad, and not a targeted one.
User goes to news.com. adserver.news.com loads ads from adserver.com . Adserver.com checks cookie from adserver.shoes.com and and generates ads server side and shows it with the news.
1. This is not possible from server-side. The way cookies work, cookies set by adserver.shoes.com are only visible from documents hosted on *.shoes.com domains. And if the tracking is done client-side, by image, iframe or some other remote call, then this is trivially detected and blocked by adblocker.
2. How did "adserver.com" knew to check for "adserver.shoes.com" cookies? there might be hunderds clients, and trying to load cookies from every domain (adserver.shoes.com, adserver.computers.com, adserver.travel-to-india.com and so on) will take way too long.
Like Google AMP.
Client site rendering can easily be blocked by just domain blocking.
If it's CSS, then no.
If it's loading an external image, then also no. Certainly no more evil than any other method of getting a user's browser to make an HTTP request anyway.
If it's tracking users then maybe. Gathering data is evil unless you have a very good reason. If you're gathering it and not actually using it then that is definitely evil. If you're gathering everything in case you need it then that is also evil. If you're gathering data that's unique to individuals that's even worse. If you're gathering data that's unique to individuals, and keeping it, and using it to build up profiles by blending it with other sources, and then selling the information that's really evil.
Just gathering browser agent strings or screen resolutions though, it's not terrible. Although I do wonder why you need CSS analytics rather than just using the server log from the request for the HTML file.
From the article:
Lots of automated traffic on the web, bots, crawlers and scrapers. So if there is a way that can remove most of the automated traffic without loading any JS, is that a win?
body:hover {
background-image: url("https://underjord.io/you-was-tracked.png");
}
[...] This has a certain elegance because it actually requires mouse interaction.That is not true. I just tried it with headless chrome and triggered :hover behavior immediately just by synthesizing mouse movement the same way I would by using it as a scraper.
You can use it to clean your analytics, but yes, it's useless for fraud prevention.
Sure, it'd clean analytics up a little bit, but that would come at the expense of being able to use a CDN/cache for my CSS. Even leaving aside implementation effort, that's a big price to get some but not all bots out of my log data.
And of course, if you share the data with a third party (that includes if you use their services to collect or process the data), you should logically assume the worst.
I would probably be okay to share all kinds of data with the websites I use, if I could reasonably trust that it will never be used to identify me personally. If you have a website, and you are curious whether your audience is mostly male or female, young or old, using computers or smartphones, I can understand the need, and would not object against my presence increasing some number in the database by one.
Well, there is this problem that if you collect too many attributes about me, and if they are all connected together (as opposed to merely increasing a few independent counters), a sufficiently large collection can be used in the future to identify me uniquely, which is the part I object against. And I have no control about how you store the information, and it is reasonable for me to assume the worst.
To make an analogy with the offline world, if you e.g. give a public lecture, you can see what kind of people are in the audience. And I would feel no desire to wear a mask for the purpose of hiding my age or gender or race or whatever from the lecturer.
But I would object against someone taking my photo, using it to identify me, and writing the information of me attending given lecture on given day into some huge dossier about me, that he would later share with other shady people, so that they can all have my incredibly detailed biography (in case of Google, including also most of my private correspondence). That is definitely evil.
It's making it personally identifiable and linking it to other things where you get in the really evil territory.
edit: autocorrect wrote "prosecution" instead of "persecution". Fixed it.
In an ideal world, it should be up to the user what they want to disclose. So, perhaps there should be no logging at all. And having loaded a page, the page should work 'offline' with no further interaction with the page or site by default. I mean, that's how simple sites appear to work. That they don't work like how you think they appear to, illustrates how technologists are selling illusions for profit.
In our world however, almost all users won't care, and the rest few won't disclose anything out of suspicion of abusing that data.
Yes, I wouldn't disclose.
But because the users don't want to give their data to you, that doesn't mean that it is ethically OK to take it without them realising.
That's the rub with privacy. There needs to be some acceptance and tolerance of an individual's decisions.
Instead though, we have engineered disclosure. In a pure sense its an act of aggression against another.
Do people put on a mask when they open a door for a stranger or when they go to the market?
This kind of knowledge is strategic for both people that want to be heard or to make a sale. Each kind of person requires a different approach. Being observant doesn't require consent.
When you are on your own home, looking a website, you think you are on your own, reading something, like a book. You do not have a sense of interacting, or of being in a public space. You would have that sense of interacting if you were on an Internet forum, or discord, etc.
A book does not report on you (and I'm talking about paper books not kindles!). But all websites monitor you. This is to say that the user is gamed. They are misled into thinking they are looking at a private thing like a book, whereas the reality is that they are being monitored. It's the biggest anti-pattern. It's absolutely without meaningful consent and it is absolutely baked in to the internet. There is no privacy online.
But, in my perfect world, we should be allowed privacy and be able to go online. Oh well, maybe in the next life!
It’s fair to “see” you there; to know my words had some impact. After all, you sought my words. Do I need to know everything about you? No. And that’s where this all breaks down.
Today’s tracking is like being a door to door salesman, and after trying to sell you something, I stalk you day and night. Actually, I’m surprised someone hasn’t tried to sue/charge Google et al for stalking.
My real concern is how we have been turned inside out by corporate technology. I mean there was a story going around a couple of years ago that Facebook knew when a couple was going to split up before the couple did!
Its too much!
I agree with the argument "If you do it to extract information from your user to which they would not consent, it’s evil."
However, we tend to get caught up in the right-now and not think through consequences. If this were widely used, browsers would implement the same sorts of privacy controls they do around 3rd party cookies, JS, etc.
This seems like a more semantic way to do tracking than many other techniques. It seems like it'd be easier for browsers to manage.
I think that first-party analytics are kind of a gray area. Third-party analytics are always evil.
Self-hosted analytics takes this out of the question.
Technically, this would all be relatively easy to block with your own user style sheet. Practically, though, a lot of non-tracking sites rely on background-image for essential functionality, so you'll see a lot of breakage. It's a dilemma.
The task of filtering out bots from server logs can get really tedious even if there's JS involved. Being able to spot humans using this technique is really quite helpful.
Edit - body:hover doesn't seem to work in Firefox, but it's trivial to work around that.
That said I won't use that in the future but it's scary how easy it is.
Don't take your personal grudge out on your users by fooling them into a false sense of security Brendan.
As technologists we want to be able to look at a technology and discern if it is good or evil. Unfortunately we don't always have enough information.
Can you give an example of how this technique could be exploited in such a way?
I do find way of voting on this matter very interesting. Parsing the logs to get the results - how amusingly nerdy!
This was a big use-case for the Practical Extraction and Reporting Language 2 decades ago... :)
Today with decent JSON logs, it's also quite fun.
I haven't built a web browser, but I built a bot and it's somewhat doable to avoid getting tracked.
A browser could feed a fake user agent and format the browser to be the correct size. After that I believe it's only IP address and cookies which are easy enough to be blocked.
It even defeats the CSS tracking mentioned. "Oh someone downloaded image 6374tracker.png, but they were from UAE and are using Firefox" and are never seen again.
My only weakness on this subject is the low level headers, anyone familiar?
Data/network/transport/session layer can be detected through TCP?
So you need to fudge a few digits somehow at your router.
Like most tools, it is up to the user, as to whether or not it's "evil."
I'm reminded of that rather silly little speech at the beginning of Dark Phoenix, where Xavier lectures Jean Gray about the uses of a pen.
If I were trying to understand users in something like A/B testing, I might use the technique, but I'd probably only do so temporarily. I'd need to make sure that the practice was outlined in the privacy policy.
In context of reality it is very nice that the author even ponders about it. This is already less evil than what we can expect on the "modern" web, however nefarious and tricky the mechanism might be. But loading a resource for tracking IPs isn't really intrinsically evil.
That being said, I think there's a line. If you use it to gather the exact same info as server log's, except with bot's filtered, that seems perfectly reasonable.
If you use it to track everything a user does on your page, that seems fairly evil.
But on reflection, it’d be more than that: anything that’s hidden (within `display: none`) and gets shown on hover… hmm. Guess it’d need to ignore `display: none` in deciding to at least fetch the resources, or suffer weird “why is this not loading the background image?” complaints.
Just like JS-based analytics I don't think most users would object heavily to someone trying to differentiate them from non-interactive bot traffic to get an idea of whether they have any real readers or just bots. But I'm absolutely certain that there are people who will not concede any analysis of the visitor's agent capabilities and behavior as legitimate. The browser sets one boundary, general public another, privacy advocates another.
I personally try to err on the side of minimal analytics. But the default idea for people creating sites tends to be that Google Analytics is "fine" and something more privacy-oriented like Plausible or Fathom are the good/ethical option. I'm not sold on that though I'm glad that there are options that are less bad. I doubt most businesses won't throw analytics out entirely but they might be willing to pick something with a better ethics profile.
Tracking is a problem when it affects your privacy. For example when too much data is collected, or when the data is handed to third parties. If you can collect just what you need, and keep that data to yourself, I really don't mind.
In this case, you measure visits anonymously without affecting the website's performance. You give me what I want without taking anything from me. I am completely fine with that.
In the real world, I'd compare it to tracking how many people enter a venue, or how many beers were sold. There is no way this could be used to tell if and when I visited that venue and bought drinks. You don't need my consent to do this.
Not respecting privacy can have bad consequences. If a country want a list of people they want to kill for various reasons such as religion, sexuality, political interests or whatever else, Facebook sells it.
Why is manipulating people bad? We are manipulated every day by various psychological factors, like seeing someone eat and getting hungry, and the government for example exerts economic pressure for unneeded items like cigarettes and alcohol with a sin tax, manipulating us to not buy them as much.
You just said it, because it's not what the user wants.
.) definitely an abuse of CSS, which is meant for visual styling
.) invisible to the user, can't even be blocked with noScript (blocking all JS), and blocking CSS would make most sites unusable or at least less than "human readable"
.) could be used to track and record a lot of information that does qualify as non-anonymous, personal data (in combination with IP address and timestamp)
.) probably illegal according to the GDPR, unless you fully inform the user and get their consent first (before loading the CSS) - and allow them to opt out.
For example check out this CCC-talk by David Kriesel - in which he demonstrates how much you can figure out just from some limited, publicly available data on a news website: https://media.ccc.de/v/33c3-7912-spiegelmining_reverse_engin...
Being able to track mouse movements and hover interactions and scroll position and time spent and other activity on the page, gives you a ton more data to mine that way than when just registering page loads.
But if a mouse enters element A at time x, then enters element B at time x+n, and then enters element C at time x+m (and so on) - you can extrapolate the path the mouse most likely must have taken - and at which speed and acceleration - to be able to get to all these positions at those time points. It's just an approximation of course, but can be surprisingly accurate.