How Bear does analytics with CSS
herman.bearblog.dev
herman.bearblog.dev
I've added an edit to the essay for clarity.
I also tested it in parallel with some other analytics platforms and it actually performed better due to the fact that adblockers are more prevalent than IP sharing per reader in this context.
One guy on twitter (no longer available) used it for mouse tracking: overlay an invisible grid of squares on the page, each with a unique background image triggered on hover. Each background image sends a specific request to the server, which interprets it!
For fun one summer, I extended that idea to create a JS-free "css only async web chat": https://github.com/kkuchta/css-only-chat
Cryptographic hashes are designed to be fast. You can do 6 billion md5 hashes in a second on an MacBook (m1 pro) via hashcat and there’s only 4 billion ipv4 addresses. So you can brute force the entire range and find the IP address. Basically reverse the hash.
And that’s true even if they used something secure like SHA-256 instead of broken MD5
I am afraid, however, that this security theater is enough to pass many laws, regulations and such on PII.
Pseudo code:
if salt.day < today():
salt = {day: today(), salt: random()}
ip_hash = sha256(ip + salt.salt)Have you read the article? That is what the author's goal seems to be.
He wants to prevent multiple requests to the same page by the same IP counted multiple times.
Cache-Control: private, max-age=86400
This prevents repeat requests for normal browsers from hitting the server.
They're just getting a list of hashes per day, and associated client info. They have no idea if the same user visit them on multiple days, because the hashes will be different.
You’ll never know enough.
Even if you hash somebody's full name, you can later answer the question "does this hash match the this specific full name". Being able to answer this question implies that the anonymisation process is reversible.
That doesn't mean that hashing is enough for pure anonymity, but used properly hashes are definitely a step above something fully reversible (like encryption with a common key).
if you have an ad ID for a person, say example@example.com, and you want to deduplicate it,
if you provide them with the names, the company that buys the data can still "blend" it with data they know, if they know how the hash was generated... and effectively get back that person's email, or IP, or phone number, or at least get a good hunch that the closest match is such and such person with uncanny certainty
de-anonymization of big data is trivial in basically every case that was written by an advertising company, instead of written by a truly privacy focused business.
if it were really a non-reversible hash, it would be evenly distributed, not predictable, and basically useless for advertising, because it wouldn't preserve locality. It needs to allow for finding duplicates... so the person you give the hash to, can abuse that fact.
What they do is pretty simple. They overwrite the data fields with the text "<Anonymized>". No hashes, no identifiers, nothing. Everything is gone. Plain and simple.
Edit:
This is a different bear:
Also, Bear claims to be GDPR compliant: https://bear.app/faq/bear-is-gdpr-compliant/
Additionally in the case of ipv6 it can be tied to a specific device more often. One cannot rely on ipv6 privacy extensions to sufficiently help there.
Whether the salt can be kept indefinitely, or is rotated regularly etc is just an implementation detail, but the key with salting hashes for analytics is that the salt never leaves the client.
As explained in the article there seems to be no salt (or rather, the current date seems to be used as a salt, but that's not a random salt and can easily be guessed for anyone who wants to say "did IP x.y.z.w visit on date yy-mm-dd?".
It's pretty easy to reason about these things if you look from the perspective of an attacker. How would you do to figure out anything about a specific person given the data? If you can't, then the data is probably OK to store.
> Whether the salt can be kept indefinitely, or is rotated regularly etc is just an implementation detail, but the key with salting hashes for analytics is that the salt never leaves the client.
I think I'm missing something.
If the salt is known to the server, then it's useless for this scenario. Because given a known salt, you can generate the hashes for every IP address + that salt very quickly. (Salting passwords works because the space for passwords is big, so rainbow tables are expensive to generate.)
If the salt is unknown to the server, i.e. generated by the client and 'never leaves the client'... then why bother with hashes? Just have the client generate a UUID directly instead of a salt.
Can you elaborate on this or link to some info elaborating what you mean? I'd like to learn about it.
> I think I'm missing something.
...
> If the salt is known to the server,
That's what you were missing yes
The reason for all this bonanza is that the ePrivacy directive requires a cookie banner, "making exceptions only for data that is "strictly necessary in order to provide a [..] service explicitly requested by the subscriber or user"*.
In the end, you only have "pinky promise" that someone isn't doing more processing on the server end, so in reality it doesn't matter much especially if the cookie lifetime is short (hours or even minutes). Actually, a cookie or other (short-lived!) client-side ID is probably better for everyone if it wasn't for the cookie banners.
That being said, the previous responder's point still stands that you can brute force the salted IPs at about a second per IP with the colocated salt. Using multiple hash iterations (e.g. 1000x; i.e. "stretching") is how you'd meaningfully increase computational complexity, but still not in a way that makes use of the general "can't be practically reversed" hash guarantees.
Kidding, of course. I don't think there's a way to track users across sessions, without storing something and requiring a 'cookie notification'. Which is kind of the point of all these laws.
If I hadn't asked for permission to send hash(ip + date) then I'd sure not ask permission if I instead stored a random salt for each new 24h and sent the hash(ip + todays_salt).
This is effectively a cookie and it's not strictly necessary if it's stats only. So I think on the server side I'd just invent some reason why it's necessary for the product itself too, and make the telemetry just an added bonus.
Presumable these are stored on the server side to identify returning visitors - so instead of storing a random number for 24 hours on the client, you now have PII stored on the server. So basically there is no way to do this that doesn't require consent.
The only way to do it is to make the information required for some necessary function, and then let the analytics piggyback on it
On the telemetry server you get e.g
Function “print” used by user 123 document 345. Using that you can do things like answering how many times an average document is printed or how many times per year an average user uses the print function.
(I don't care about this Bear analytics thing at all, and just clicked the comment thread to see if it was the Bear I thought it was; I do care about people's misconceptions about hashing.)
The only thing you can brute force from that is some IP and some salt such that SHA1(IP+Salt) = 6FF6BA399B75F5698CEEDB2B1716C46D12C28DF5 But you'll find millions of such IPs. Perhaps all possible IP's will work with some salt, and give that hash. It's not revealing my IP even if you manage to find a match?
Yes if all you ever want to send is a unique visitor ID then there is no point in having a local hash, because you can just generate a random ID and use that to identify the user.
What I mean is that if you want to send multiple pieces of PII (such as an IP, a filename, a username,...) then the only way to do that safely is to send hash(salt+filename) for example, where the salt is not known to the server receiving the hash. The IP in the suggestion to use a locally stored hash here just represented "PII that should be sent anonymously" and not "A good way of identifying a unique system".
Sure, but it does at least prevent the use of rainbow tables. Arguably not relevant in this scenario, but it doesn't mean that salting does nothing. Rainbow tables can speed up attacks by many orders of magnitude. Salting may not prevent each individual password from being brute forced, but for most attackers, it probably will prevent your entire database from being compromised due to the amount of computation required.
"Salted hashes" are one of the more treacherous security cargo cults, because they create the impression that the big problem you have to solve hashing a password is somehow mixing in a salt. No: just doing that by itself doesn't accomplish anything meaningful at all.
Not really. They are designed to be fast enough and even then only as a secondary priority.
> You can do 6 billion … hashes/second on [commodity hardware] … there’s only 4 billion ipv4 addresses. So you can brute force the entire range
This is harder if you use a salt not known to the attacker. Per-entry salts can help even more, though that isn't relevant to IPv4 addresses in a web/app analytics context because after the attempt at anonymisation you want to still be able to tell that two addresses were the same.
> And that’s true even if they used something secure like SHA-256 instead of broken MD5
Relying purely on the computation complexity of one hash operation, even one not yet broken, is not safe given how easy temporary access to mass CPU/GPU power is these days. This can be mitigated somewhat by running many rounds of the hash with a non-global salt – which is what good key derivation processes do for instance. Of course you need to increase the number of rounds over time to keep up with the rate of growth in processing availability, to keep undoing your hash more hassle than it is worth.
But yeah, a single unsalted hash (or a hash with a salt the attacker knows) on IP address is not going to stop anyone who wants to work out what that address is.
That's not a reasonable way to say it. It's literally the second priority, and heavily evaluated when deciding what algorithms to take.
> This is harder if you use a salt not known to the attacker.
The "attacker" here is the sever owner. So if you use a random salt and throw it away, you are good, anything resembling the way people use salt on practice is not fine.
For a bit of clarity around IP addresses hashes. The only use they have in this context is preventing duplicate hits in a day (making each page view unique by default). At the end of each day there is a worker job that scrubs the ip hash which is now irrelevant.
Yes, these are marginal groups perhaps, but it is always super bad sign seeing them excluded in any way.
I am not sure (I doubt) there is a 100 % reliable way to detect that "real user is reading this article (and issue HTTP request)" from baseline CSS in every single user agent out there (some of them might not support CSS at all, after all, or have loading of any kind of decorative images from CSS disabled).
There are modern selectors that could help, like :root:focus-within (requiring that user would actually focus something interactive there, what again is not guaranteed for al agents to trigger such selector), and/or bleeding edge scroll-linked animations (`@scroll-timeline`). But again, braille readers will probably remain left out.
Joking aside, I love to read websites with keybaords, esp. if I'm reading blogs. So, it's possible that sometimes my pointer is out there somewhere to prevent distraction.
[1] was it base ten, right?
OP just thought of a creative, effective and probably faster more code efficient way to do analytics. I love it, thanks OP for sharing it
> Note: The :hover pseudo-class is problematic on touchscreens. Depending on the browser, the :hover pseudo-class might never match, match only for a moment after touching an element, or continue to match even after the user has stopped touching and until the user touches another element. Web developers should make sure that content is accessible on devices with limited or non-existent hovering capabilities.
I have to have two interfaces for my web apps, because sometimes I want to hide some text of marginal value behind a :hover, but only a mouse can see it unless I break it out somehow for touch.
No, the reason is just because it’s too fiddly and too unreliable (to use, I mean, more than the touchscreen itself being unreliable, though the more sensitive you tune it the more that will become a problem as well).
Apple had their 3D Touch and they dropped it because (from what I hear, I never used it) it largely confused people due to non-obviousness/non-familiarity, required more care to use accurately than people liked (or, in some cases, were able to provide), and could become unreliable.
I’ve used a variety of “normal” capacitive touchscreens that could be activated anywhere from a few millimetres from the surface to requiring a firm/large-surface-area touch. It’s all about how they’re calibrated, and how much they’ve deteriorated (public space ones seem to often go very bad astonishingly quickly).
Most recently, in Indian airports the checkin machines that you can choose to use say that they’re touchless, that you can just put your finger near and it will work and isn’t that all lovely and COVID-19-aware of them, but on three separate machines I’ve tried it very carefully and couldn’t get it to activate before touching the screen.
I believe that's as they're trying to live in what amounts to a toxic wasteland. Users like us are done with the whole concept and as such I assume if CSS analytics becomes popular, then attempts will be made to bypass that too.
||somesite.example^$css
would work in ublockI manually unblocked Piwik/Matomo, Plausible and and Fathom from ublock. I don't see any harm in what and how these track. And they do give the people behind the site valuable information "to improve the service".
e.g. Plausible collects less information on me than the common nginx or Apache logs do. For me, as blogger, it's important to see when a post gets on HN, is linked from somewhere and what kinds of content are valued and which are ignored. So that I can blog about stuff you actually want to read and spread it through channels so that you are actually aware of it.
And also just knowing some people read what you write is nice. There is nothing wrong with having some validation (as long as you don't obsess over it) and it's a basic human need.
This is just for a blog; for a product knowing "how many people actually use this?" is useful. I suspect that for some things the number is literally 0, but it can be hard to know for sure.
User interviews are great, but it's time-consuming to do well and especially for small teams this is not always doable. It's also hard to catch things that are useful for just a small fraction of your users. i.e. "it's useful for 5%" means you need to do a lot of user interviews (and hope they don't forget to mention it!)
Services like Plausible give you the bare minimum to understand what is viewed most. If you have a website that you want people to visit then it’s a pretty basic requirement that you’ll want to see what people are interested in.
When you start “personalising” the experience based on some tracking that’s when it becomes a problem.
not really
it should be what you are competent and proficient at
people will come because they like what you do, not because you do the things they like (sounds like the same thing, but it isn't)
there are many proxies to know what they like if you want to plan what to publish and when and for how long, website visits are one of the less interesting.
a lot of websites such as this one get a lot of visits that drive no revenue at all.
OTOH there are websites who receive a small amount of visits, but make revenues based on the amount of people subscribing to the content (the textbook example is OF, people there can get from a handful of subscriber what others earn from hundreds of thousands of views on YT or the like)
so basically monitoring your revenues works better than constantly optimizing for views, in the latter case you are optimizing for the wrong thing
I know a lot of people who sell online that do not use analytics at all, except for coarse grained ones like number of subscriptions/number of items sold/how many email they receive about something they published or messages from social platforms etc.
that's been true in my experience through almost 30 years of interacting and helping publishing creative content online and offline (books, records, etc)
This isn’t true for all channels. The current state of search requires you to adapt your content to what people are looking for. Social channels are as you’ve said.
It doesn’t matter how you want to slice it. Understanding how many people are coming to your website, from where and what they’re looking at is valuable.
I agree the “end metric” is whatever actually drives the revenue. But number of people coming to a website can help tune that.
you chose to send an email or to buy a product or a subscription, which is different from being tracked.
it's still a metric, but has an higher value.
it's people genuinely interested in what you offer.
If every web client stopped the tracking, you, as blogger, could go back to just getting analytics on server logs (real analytics, using maths).
Arguably state of the art in that approach to user/session/visits tracking 20 years ago beats today's semi-adblocked disaster. By good use of path aliases aka routes, and canonical URLs, you can even do campaign measurement without messing up SEO (see Amazon.com URLs).
||bearblog.dev/hit/
This is the shortest it can be written with certainty of no false positives, but you can do things like making the URL pattern more specific (e.g. /hit/*/) or adding the image option (append $image) or just removing the ||bearblog.dev domain filter if it spread to other domains as well (there probably aren’t enough false positives to worry about).I find it also worth noting that all of these techniques are pretty easily circumventable by technical means, by blending content and tracking/ads/whatever. In case of all-out war, content blockers will lose. It’s just that no one has seen fit to escalate that far (and in some cases there are legal limitations, potentially on both sides of the fight).
The Chrome Manifest v3 and Web Environment Integrity proposals are arguably some of the clearest steps in that direction, a long term strategy being slow-played to limit pushback.
I'm not sure how many people actually fear that. Might get responses from "yes, and it's creepy" to "don't be daft that's just SciFi".
If anything, you're gonna get numbers that are inflated because it's a bit impossible to dismiss all of the bot traffic just by looking at user agents.
IPs are useful in case of attack, but you could limit yourself to simply logging subnets. It's a little more aggressive block a subnet, or an entire ISP, but it seems like a good tradeoff.
The primary reason I care about analytics is to see if posts are getting read, which on the surface (and in some ways) is for reasons of vanity, but is actually about writer-reader engagement. I'm genuinely interested in what my readers resonate with, because I want to give them more of that. The "that" could be topical, tonal, length, who knows. It helps me hone my material specifically for my readers. Ultimately, I could write about a dozen different things in two dozen different ways. Obviously, I do what I like, but I refine it to resonate with my audience.
In this sense, analytics are kind of a way for me to get to know my audience. With blogs that had high engagement, analytics gave me a sort of fuzzy character description of who my readers were. As with above, I got to see what they liked, but also when they liked it. Were they reading first thing in the morning? Were they lunch time readers? Were they late at night readers. This helped me choose (or feel better about) posting at certain times. Of course, all of this was fuzzy intel, but I found it really helped me engage with my readership more actively.
You might get no monetary value from having 12 people read the site or 12,000 but from a personal perspective it's nice to know what people want to read about from you, and so you can feel like the time you spent writing it was well spent, and adjust if you wish to things that are more popular.
And the real benefit of this trick is separating users from bots.
That you cannot access the logs because you don't own the servers doesn't mean there aren't any servers that have logs.
No. The most obvious way is to reassess running on servers/services that you don't own, and which don't offer features you need.
This is explained in the blog post:
> There's always the option of just parsing server logs, which gives a rough indication of the kinds of traffic accessing the server. Unfortunately all server traffic is generally seen as equal. Technically bots "should" have a user-agent that identifies them as a bot, but few identify that since they're trying to scrape information as a "person" using a browser. In essence, just using server logs for analytics gives a skewed perspective to traffic since a lot of it are search-engine crawlers and scrapers (and now GPT-based parsers).
Let's say I have an e-commerce website, with products I want to sell. In addition to analytics, I decide to log a select few actions myself such as visits to product detail page while logged in. So I want to store things like user id, product id, timestamp, etc.
How do I actually store this? My naive approach is to stick it in a table. The DBA yelled at me and asked how long I need data. I said at least a month. They said ok and I think they moved all older data to a different table (set up a job for it?)
How do real people store these logs? How long do you keep them?
I once did this, and we didn’t need to even think about partitioning until we hit a billion rows or so. (But partition sooner than that, it wasn’t a pleasant experience)
They can do aggregations much faster and can deal with sparse/many columns (the "paid" event has an "amount" attribute, the "page_view" event has an "url" attribute...)
I am approaching 1.5 million rows in under two months. Thankfully, my DBA is kind, generous, and infinitely patient.
Clickhouse looks like a good approach. I'll have to look into that.
> select count(*) from trackproductview;
> 1498745
> select top 1 createddate from TrackProductView order by createddate asc;
> 2023-08-18 11:31:04.000
what is the maximum number of rows in clickhouse table? Is there such a limit?
It’s all a heuristic and even with high collision hashing, analytics would provide some additional insight.
Analytics is not about absolute accuracy, it's about measuring differences; things like which pages are most popular, did traffic grow when you ran a PR campaign etc.
> ‘personal data’ means any information relating to an identified or identifiable natural person (‘data subject’); an identifiable natural person is one who can be identified, directly or indirectly, in particular by reference to an identifier such as a name, an identification number, location data, an online identifier or to one or more factors specific to the physical, physiological, genetic, mental, economic, cultural or social identity of that natural person;
In this context, if I define a hashing function that e.g. sums all ip address octets, what then?
Summing octets is non-reversable, so it seems like a good 'hash' to me (but note: you'll get a lot of collisions). And of course, IANAL.
The linked article talks about identification numbers that can be used to link a person. I am not a lawyer but the article specifically refers to one person.
By that logic, if the hash you generate cannot be linked to exactly one, specific person/request - you’re in the clear. I think ;)
Having that IP and user timezone you can generate the same hash and trace back the user. This is hardly anonymous hashing.
Wide Angle Analytics adds daily, transient salt to each IP hash which is never logged thus generating a truly anonymous hash that prevents reidentification.
If it’s so destructive that it’s impossible to track users, it’s useless for you. If not, you need a privacy banner.
IMO (and I'm not a lawyer): if you store ip+site for 24 hours and after that only store "region" (maybe country or state) and site this should be GDPR compliant.
[0] it should use sha256 or similar and not md5
Not a layer but discussed this previously with lawyers when building a GDPR framework awhile back.
I've been told this directly by lawyers who specialize in GDPR and CCPA etc. I will take their word over yours.
If you are a lawyer with direct expertise in this area then I'm willing to listen.
My scraping bots use an instance of chrome and therefore trigger hover as well, but you'll cut out the less sophisticated bots.
This is because of protection systems, if I try to scrape my target website with just code I just get insta banned / "captched".
You totally could be triggering it, but not every bot will, even the fancy ones.
But maybe I'm not aware that browsers guarantee that "images" loaded using url() in CSS will be (re)loaded exactly once per page?
This bit me when I tried to make a page that reload an image as a form of monitoring. However URL interestingly includes the fragment (after the #) even though it isn't set to the server. So I managed to work around this by appending #1, #2, #3... to the image URL.
https://gitlab.com/kevincox/image-monitor/-/blob/e916fcf2f9a...
Many ISPs are now using CG-NAT so this approach would miscount thousands of visitors seemingly coming from a single IP address.
(It would be better if he used a hash of the raw user agent string)
* You cannot (easily) track interaction events (esp. relevant for SPAs, but also things like "user highlighted x" or "user typed Y, then backspaced then typed Z)"
* You cannot track timings between events (e.g. how long a user is on the page)
* You cannot track data such as screen-sizes, agents, etc.
* You cannot track errors and exceptions.
re the hashing issue, it looks interesting but adding more entropy with other client headers and using a stronger hash algo should be fine.
I can imagine many cases where real human user doesn't scroll the page on mobile platform. I like the CSS approach but I'm not sure it's better than doing some bot filtering with the server logs.
I think we would either need to send fake data to these analytics tools deliberately like https://adnauseam.io/
Or now include CSS as a spy tracker that needs to be blocked.
For this, this would need to block the endpoint or send obfuscated data deliberately in protest of this.
Should you want to cover access logs also, then forms of tracking then sending excessive, random obfuscation data with adnauseam would also help here.
If not, what's wrong with a service knowing you're accessing it? How can they serve a page without knowing you're getting a page?
||/hit/*$image
In your favorite ad blocker- I'd decouple the database writes with a message queue.
- Make sure that table is partitioned if you haven't. TimescaleDB's hyper tables are a good compromise here if you are already using Postgres..
> Now, when a person hovers their cursor over the page (or scrolls on mobile) it triggers body:hover which calls the URL for the post hit
> The :hover pseudo-class is problematic on touchscreens. Depending on the browser, the :hover pseudo-class might never match
https://developer.mozilla.org/en-US/docs/Web/CSS/:hover
Don't take my word for it. Trying it in mobile emulators will have the same result.
The basic server logs include declared bots, undeclared bots pretending to use browsers, undeclared bots actually using browsers, and humans.
The hit endpoint logs will exclude almost all declared bots, almost all undeclared bots pretending to use browsers, and some humans, but will retain a few undeclared bots that search for and load subresources, and almost all humans. About undeclared bots that actually use browsers, I’m uncertain as I haven’t inspected how they are typically driven and what their initial mouse cursor state is: if it’s placed within the document it’ll trigger, but if it’s not controlled it’ll probably be outside the document. (Edit: actually, I hadn’t considered that bearblog caps the body element’s width and uses margin, so if the mouse cursor is not in the main column it won’t trigger. My feeling is that this will get rid of almost all undeclared bots using browsers, but significantly undercount users with large screens.)
But my experience is that reasonably simple heuristics do a pretty good job of filtering out the bots the hit endpoint also excludes.
• Declared bots: the filtration technique can be ported as-is.
• Undeclared bots pretending to use browsers: that’s a behavioural matter, but when I did a little probing of this some years ago, I found that a great many of them were using unrealistic user-agent strings, either visibly wonky or impossible or just corresponding to browsers more than a year old (which almost no real users are using). I suspect you could get rid of the vast majority of them reasonably easily, though it might require occasional maintenance (you could do things like estimate the browser’s age based on their current version number and release cadence, with the caveat that it may slowly drift and should be checked every few years) and will certainly exclude a very few humans.
• Undeclared bots actually using browsers: this depends on the unknown I declared, whether they position their mice in the document area. But my suspicion is that these simply aren’t worth worrying about because they’re not enough to notably skew things. Actually using browsers is expensive, people avoid it where possible.
And on the matter of humans, it’s worth clarifying that the hit endpoint is worse in some ways, and honestly quite risky:
• Some humans will use environments that can’t trigger the extra hit request (e.g. text-mode browsers, or using some service that fetches and presents content in a different way);
• Some humans will behave in ways that don’t trigger the extra hit request (e.g. keyboard-only with no mouse movement, or loading then going offline);
• Some humans will block the extra hit request; and if you upset the wrong people or potentially even become too popular, it’ll make its way into a popular content blocker list and significant fractions of your human base will block it. This, in my opinion, is the biggest risk.
• There’s also the risk that at some point browsers might prefetch such resources to minimise the privacy leak. (Some email clients have done this at times, and browsers have wrestled from time to time with related privacy leaks, which have led to the hobbling of what properties :visited can affect, and other mitigations of clickjacking. I think it conceivable that such a thing could be changed, though I doubt it will happen and there would be plenty of notice if it ever did.)
But there’s a deeper question to it: if you don’t exclude some bots; or if the URL pattern gets on a popular content filter list: does it matter? Does it skew the ratios of your results significantly? (Absolute numbers have never been particularly meaningful or comparable between services or sources: you can only meaningfully compare numbers from within a source.) My feeling is that after filtering out most of the bots in fairly straightforward ways, the data that remains is likely to be of similar enough quality to the hit endpoint technique: both will be overcounting in some areas and undercounting in others, but I expect both to be Good Enough, at which point I prefer the simplicity of not having a separate endpoint.
(I think I’ve presented a fairly balanced view of the facts and the risks of both approaches, and invite correction in any point. Understand also that I’ve never tried doing this kind of analysis in any detail, and what examination and such I have done was almost all 5–8 years ago, so there’s a distinct possibility that my feelings are just way off base.)
I use it for my personal stuff and it's great. No hassle, just paste your markdown in and you're good to go.
The simple solution is to respect the basic wishes of those who do not want to be tracked. This is a "struggle" only because website operators don't want to hear no.
It's usually website hosts just wanting to know how many folks are passing through. If a visitor doesn't even want to contribute to incrementing a private visit counter by +1, then maybe don't bother visiting.
Why aren't people starting their protest against the website they're currently using instead of at OP.
You're upset that OP used an image or javascript instead of grepping `access.log` makes absolute no sense. The same data is shown there.
One difference is intent. When you build an analytics system you have an obligation to let people opt out. Access logs serve many legitimate purposes, and yes, they can also be used to track people, but that is not why access logs exist. This difference is also reflected in law. Using access logs for security purposes is always allowed but using that same data for marketing purposes may require an opt-in or disclosure.
That doesn't really work because huge amount of traffic is from 1) bots, 2) prefetches and other things that shouldn't be counted, 3) the same person loading the page 5 times, visiting every page on the site, etc. In short, these numbers will be wildly wrong (and in my experience "how wrong" can also differ quite a bit per site and over time, depending on factors that are not very obvious).
What people want is a simple break-down of useful things like which entry pages are used, where people came from (as in: "did my post get on the frontpage of HN?")
I don't see how anyone's privacy or anything is violated with that. You can object to that of course. You can also object to people wearing a red shirt or a baseball cap. At some point objections become unreasonable.
Tons of very simple hosts, like Github Pages, don't give access to detailed info like that (and that's totally fine).
For example, if you walked into my coffee shop, I would be able to lay eyes on you and count your visits for the week. I could also observe were you sit and how long you stay. If I were to better serve you with these data points, by reserving your table before you arrive with your order ready, you'd probably welcome my attention to detail. However, if I were to see you pulled about x number of watts a month from my outlets, then locked up the outlets for a fee suddenly - then you'd rightfully wish to never be observed again.
So what I'm getting at is, the issues with tracking appear to be with the perverse assholes vs. the benevolent shopkeeps of the tracking.
To wrap up this thought: what's happening now though is a stalker is following us into every store, watching our every move. In physical space, we'd have this person arrested and assigned a restraining order with severe consequences. However, instead of holding those creeps accountable, we've punished the small businesses that just want to serve us.
--
I don't know how I feel about this or really what to do.
My point is that I don't get why the small businesses should have the right to track me to offer me better services that I never even asked for. Sure, its nice, but its not worth deregulating tracking and allowing all the evil corps to track me too.
You walk into your favorite coffee shop, order your favorite coffee, every day. But because of privacy reasons the coffeeshop owner is unaware of anything. Doesn't even track inventory, just orders whatever whenever.
One day you walk in and now you can't get your favorite coffee... Because the owner decided to remove that item from the menu. You get mad, "Where's my favorite coffee?" the barista says "owner removed it from menu" and you get even more upset "Why? Don't you know I come in here every day and order the same thing?!"
Nope, because you don't want any amount of tracking whatesoever, knowing any type of habits from visitors is wrong!
But in this scenario you deem the owner knowing that you order that coffee every day ensures that it never leaves the menu, so you actually do like tracking.
Assuming you live in the US, next time you're in a grocery store, count how many cameras you can spot. Then consider: these cameras could possibly have facial recognition software; these cameras could possibly have object recognition software; these cameras could possibly have software that tracks eye movements to see where people are looking.
Then wonder: do they have cameras in the parking lot? Maybe those cameras can read license plates to know which vehicles are coming and going. Any time I see any sort of news about information that can be retrieved from a photo, I assume that it will be done by many cameras at >1 Hertz in a handful of years.
Current website tracking is like the coffee shop owner hiring a private investigate to dig into the personal lives of everyone who walks in the door so they can suggest the right coffee and custom cup without having to ask. They could not do that and just let someone pick their own cup... or give them a generic one. I'd like that better. If clipboards in coffee shops are dystopian, so is current web tracking, and we should feel the same about it.
I think Bear strikes a good balance. It lets authors know someone is reading, but it's not keeping profiles on users to target them with manipulative advertising or some kind of curated reading list.
<cynical_statement> The purpose of the web is to distribute ads. The "struggle" is with people who think we made this infrastructure so you could share recipes with your grand-mother. </cynical_statement>
The interstate highway system in the US wasn't build with the intent of advertising to people, it was to move people and goods around (and maybe provide a means to move around the military on the ground when needed). Once there were a lot of eyes flying down the interstate, the billboard was used to advertise to those people.
The same thing happened with the newspaper, magazines, radio, TV, and YouTube. The technology comes first and the ads come with popularity and as a means to keep it low cost. We're seeing that now with Netflix as well. I'm actually a little surprised that books don't come with ads throughout them... maybe the longevity of the medium makes ads impractical.
But JS trackers are so much more. Time spent on the website, scroll depth, screen sizes, some limited and compliant and yet useful unique sessions, those things cannot be achieved without some (simple) JS.
Server side, JS, CSS... No one size fits all.
Wide Angle Analytics has strong privacy, DNT support, an opt-out mechanism, EU cloud, compliance documentation, and full process adherence. Employs non-reversible short-lived sessions that still give you good tracking. Combine it with custom domain or first-party API calls and you get near 100% data accuracy.
The Cloud Act rendered that worthless.
The CSS tracker is as useful as server log-based analytics.
It is not. Have you read the article?The whole point of the CSS approach is to weed out user agents which are not doing mouse hover on the body events. You can't see that from server logs.
(not the creator, just a regular user of this great tool)
We need more tools that send random, fake data to analytics providers which renders the analytics useless to them in protest of tracking.
If there are any more like Adnauseam, I would love to know.