Show HN: Trackless - A GDPR-Friendly Google Analytics Opt-In Button
github.com
github.com
[...]
> You must ask people to actively opt in. Don’t use pre-ticked boxes, opt-out boxes or other default settings.
Source:
https://ico.org.uk/for-organisations/guide-to-the-general-da...
Anonymous data is specifically excluded from GDPR. Google Analytics provides an IP anonymization feature. If you're absolutely confident that your users can't be personally identified based on the data being sent to Google Analytics, then you don't need consent.
https://gdpr-info.eu/recitals/no-26/
https://support.google.com/analytics/answer/2763052?hl=en
https://support.google.com/analytics/answer/6366371?hl=en&re...
The moment you load a resource on your page from an external source, you lose almost all control of what the operator of that external source does with any personal data that your visitor's browser sends to them, any cookies it sends with its reply, or what it does more generally in the case of executable resources.
Given that modern web sites routinely incorporate external assets for a multitude of reasons, has anyone ever found any official, authoritative guidance on who is the data controller or data processor in such cases, how they are expected to meet any obligations they have in terms of transparency and obtaining consent, or the related question of who is responsible for giving notifications or obtaining consent if required under the "cookie law"?
But that's not what how the controller is defined in the regulations. To be the controller, you must be "the natural or legal person, public authority, agency or other body which, alone or jointly with others, determines the purposes and means of the processing of personal data". If you don't even have any way to know what personal data a third party is collecting or how it's being used, and you're linking to content that is freely available but over which you have no control, you're not even close to fitting that definition.
You should have agreements in place with all your external resource providers that touch personal data.
But that fundamentally breaks most of the modern WWW, which is not a reasonable thing to do. You can't even have a personal blog linking to a jQuery CDN to expand or contract your sidebar or Google Fonts to make things look pretty at that point.
Just because you are using a third party to do the spying does not remove your responsibility.
If they're spying at your request and on your behalf, that's one thing.
But it is inherent in the technologies of the web that third parties may be doing all kinds of things without your knowledge, consent or control. Moreover, even if you have somehow satisfied yourself that there is nothing inappropriate going on when you first incorporate external content in your page by reference, there is in general no technical mechanism to guarantee that the situation will not change later. In some limited cases tools like subresource integrity can help, but they only address specific parts of the general issue.
"‘controller’ means the natural or legal person, public authority, agency or other body which, alone or jointly with others, determines the purposes and means of the processing of personal data"
If you're embedding a JS library from a CDN, then you have a lawful basis for passing the IP address of your user to a third party under Art. 6(1)(f). As long as you've performed a reasonable risk assessment about this activity and have records to prove it, you should satisfy your obligations as a controller under Chapter 4. If they go rogue and add a bunch of tracking scripts to the library, they're liable. You'd still need to notify about the breach.
If you're embedding a Javascript ad unit that does a bunch of tracking, you probably don't have a legitimate interest under Art. 6(1)(f), so you'll need consent. You're intentionally passing a bunch of personal data to a third party, so your responsibilities with regards to risk assessment are far greater. You and the ad provider probably constitute joint controllers under Art. 26.
https://gdpr-info.eu/recitals/no-15/
https://gdpr-info.eu/art-6-gdpr/
https://gdpr-info.eu/chapter-4/
https://gdpr-info.eu/art-26-gdpr/
IANAL etc.
But if you're embedding a JS library from a CDN, then as a matter of fact, you aren't passing any data about your user to the third party at all. The user's browser is doing that as part of its normal operation.
Moreover, as another matter of fact, you cannot have either any knowledge or any control over what happens next regarding any personal data the third party is collecting or how it is being processed, unless you have some separate arrangement with the third party that goes well beyond mere linking or embedding.
Logically, it doesn't seem to make much sense for you to be either the controller or the processor in that instance. However, if the third party plays either role, they may have no mechanism to communicate with your site visitor to fulfil their obligations either.
You're not absolutely and totally responsible for anything that might possibly happen under any circumstances, but you're required to implement appropriate technical and organisational measures to ensure and to be able to demonstrate that processing is performed in accordance with this Regulation (Art. 24).
If you embed a JS library from Google Cloud CDN and compile a written risk assessment with copies of Google's privacy policy for that product and their EU-US Privacy Shield and ISO 27018 certifications, you're probably fine no matter what happens.
If you embed a JS library from SuspiciousCDN.ru because someone on 4chan gave you the link and your user data ends up on WikiLeaks, you're going to have some serious explaining to do.
Do I expect visitors to my personal blog to have any security or privacy problems because I use Google Fonts to make it look nice? No.
Do I have any sort of formal agreement that is binding on Google to guarantee that, or that obliges them to notify me if they change their privacy policy in this respect? Also no.
And exactly the same applies to, for example, numerous popular JS libraries that are hosted for free on reputable CDNs.
It really all does revolve around control: as the website owner you can't control their ISP, Browser, VPNusage, etc. but you could trivially change your site from using Google Analytics to using a different system (or disable it altogether).
But the regulations don't say anything about websites. Being a controller is about whether you determine the purposes and means of processing (even if someone else is then doing that processing). Merely embedding third party content on your site doesn't even give you knowledge of that processing, never mind any control over it.
Now, if you have a formal agreement with some other service that they will process personal data for some purpose and you will embed something in your page that gives them access to that data, obviously in that situation you're acting as controller. But huge amounts of the embedding that happens in the real world don't have those formal arrangements, and if you say no-one can ever embed anything any more without formal legal agreements to protect themselves, you break a large part of the modern web both technically and culturally. I don't think that was the intent of the new regulations, nor do I think it is a reasonable thing to do.
Take it as, "I control the door to a bank vault, if I allow robbers in, I will be a complice to a crime as the crime couldn't be commited without your help". Negligence or direct intent, it can be costly. Assess your 3rd party sources very carefully, I have already removed GA and replaced them with local analytics (https://matomo.org/) as I can't trust them, they are trying to downplay GDPR and there is already a complaint written against them (https://noyb.eu not for GA though), and I have read the PDFs, they are right and quite objectively, they are guilty. I dont want to be in a same boat with them.
Yes there is a guidance, it is called GDPR, it is THE only guidance, just take the concepts, I can give you this link, it is the best I was able to find, it will help understand the GDPR, but for each and every site, owner needs to decide on its own: https://www.youtube.com/watch?v=-stjktAu-7k
The modern web depends on embedding third party content for many reasons, most of which have nothing to do with invading anyone's privacy and many of which are directly in the visitor's interests. It is not helpful to undermine that whole ecosystem and expect everyone to start having formal contracts in place before they can take advantage of any of those services. Nor is it reasonable to expect services offered for free that aren't doing anything shady to take on significant liability and/or other commitments anyway through formal agreements with their users. Why would they do that, instead of just (as obviously quite a few places already have) geoblocking the EU to remove themselves from the scope of the onerous rules?
To the morons (no, it is not insult, it is empirical fact) downvoting me, it is not me, it is GDPR, face the reality, it is not my fault that you are too reluctant to understand it and biting people trying to help you out wont help. Downvoting me wont change GDPR or change anything, you will just loose a valuable source of information as you did just now. Go to the first psychiatrist and it will tell you that a reality will be as it is even if you close your eyes (or shoot the messenger =/).
Don't forget to upvote me, when you figure out I was right and you get a warning/fine.
If you don't want to be banned, you're welcome to email hn@ycombinator.com and give us reason to believe that you'll follow the rules in the future.
The &aip=1 feature - in spite of it's name - does not provide any useful anonymity! As you can see in Google's own documentation (your 2nd link), when aip=1 GA claims that "the last octet of the user IP address is set to zero".
At best this can only group your IP with the neighboring 255 addresses. Google still logs the upper 24-bits of the address, which is probably enough to discover e.g. your ASN and geolocation. In practice, IP addresses usage is not perfectly uniform, so your actual "anonymity" is less than the theoretical maximum of 1-in-256. In general, the HTTP headers, cookies, etc will have at least 8 bits of unique entropy that more than makes up for losing the least interesting 8 bits of your IPv4 address.
This feature isn't designed to provide actual anonymity. The documentation even suggests the feature was designed to minimally satisfy certain legal or contractual obligations:
>> This feature is designed to help site owners comply with their own privacy policies or, in some countries, recommendations from local data protection authorities, which may prevent the storage of full IP address information.
Notice that this mentions pre-GDPR "recommendations" and that compliance is the goal, not user anonymity.
(side note: that documentation doesn't even acknowledge IPv6. Does the aip=1 feature even exist for IPv6?)
Recital 26:
The principles of data protection should apply to any information concerning an identified or identifiable natural person... To determine whether a natural person is identifiable, account should be taken of all the means reasonably likely to be used, such as singling out, either by the controller or by another person to identify the natural person directly or indirectly. To ascertain whether means are reasonably likely to be used to identify the natural person, account should be taken of all objective factors, such as the costs of and the amount of time required for identification, taking into consideration the available technology at the time of the processing and technological developments.
If I give you an IP with the last octet redacted, how would you use that information to identify a natural person? If you can think of a method, how long does it take? How much does it cost? Is it reasonably likely to be used?
That depends a lot on 1) the other data that submitted in the same set of analytics events. and 2) the data found in other databases that might correlate with the data in #1.
> how long does it take?
How long does it take to run a SELECT statement that JOINs a handful of large tables? This could be any amount of time, but I suspect anybody with a lot of resources like Google can probably run this kind of query (e.g. map all analytics records to personal gmail accounts) ad-hoc in minutes. A better idea would be to integrate the correlation into the handling of analytics events.
> How much does it cost?
How much does it cost to run a large query on your DB? The only real expenses would derive from the volume of analytics events want to process per second. Mapping a single analytics event to existing databases would be approximately free.
> Is it reasonably likely to be used?
I have very little doubt that at least Google and FB do this kind of re-correlation in some situations. I have no how common the practice would be.
--
These questions suggest you might be missing just how trivial this problem is to solve. Google already has massive databases that identify a "natural person" (like a gmail account associated with a mobile telephone number for 2FA). Unrelated to GA, the databases handling regular gmail activity can store [IP addr, other TCP/IP headers, HTTP headers, accurate (~1s) timestamps] simply because your browser made a HTTP request over a TCP socket to fetch the text of your email.
With all those resources available, Google receives a GA event, notices aip=1, and dutifully sets the least significant 8 bits to 0. At that point they simply use the other 24 bits to search the recent logs for matching HTTP requests. This may already select a unique account, but in general it -probably selects about 200 to 500. (256 from the ambiguity of not using 8 bits of address, multi0plied by the average number of gmail users behind the same NATed address)
That was the easy part, which defines the real problem as finding the real account out of a selection of a few hundred. So start trying to correlate the rest of the available data. Did the GA event contain a UserAgent string that is unique with respect the few hundred in our search space? If that wasn't unique, repeat with every other HTTP header. If still not unique, try longer tuples where the entire tuple must match. Repeat for any other available data.
I could get into the interesting ways you could exploit non-random IP numbers (how does your router rewrite TCP Source Port? Do your TCP Initial Sequence Numbers reveal your OS?[1]), but that level of analysis probably isn't necessary. An important question at thi8s point is how much error is acceptable? Even if the previous searches did not result in a unique match, they probably reduced the search space down to only a handful of candidates. Start apply Bayes Theorem[3] or other statistical analysis methods; is there a match with an acceptable confidence? What about a larger network[4] of inferences?
There are many ways to approach the problem of finding the correct record out of a few hundred; I'm only sketching a fairly straightforward method. I'm sure Google and FB can do fancier things with better techniques such as machine learning. The point is that 24 bits of identifying entropy is a lot. It's already so close to being a unique identifier, constructi8ng an actual unique ID only requires adding a few bits of entropy, which probably available in the surrounding metadata and/or session data.
[1] The ISN shouldn't reveal much in modern OS. However, reading this[2] paper about how they used to be broken was really enlightening when I read it when it was originally published. The visuals demonstrate clearly how easily your difficult/random searches can collapse into a trivial search space.
[2] http://lcamtuf.coredump.cx/oldtcp/tcpseq/print.html
Check my post below, I would be glad if you have some idea, but as far as I am concerned, anonymising IP to keep getting uniform result is tehnically impossible.
I am asking this as a friend of mine is having hard time accomplishing exactly that and is really a hard nut to crack, anonymization is by default irreversable and making such algorythm for 4 numbers (actually even less due to known ip address ranges for EU users + reserved ranges) is not simple. You can seed it but that key must remain unknown to google, while this is again getting very hard with javascript. The only way I see is sending all the data to local proxy script, anonymizing the data on your side and then sending it to GA.
I thing that if GA is doing just some hashing, this opens all the sites, using it, to a GDPR responsibility as data controllers including HN. And this can't be hidden under capet (imho) as a "I can't offer service without it" (legitimate interest).
If you enable the Anonymize setting, the last octet (IPv4) or last 80 bits (IPv6) is set to zero by the analytics collector. The full IP is never stored or processed.
Yes, I see why your marketing sense told you to choose the untruthful version.
Watch out with GDPR, this is not cookie law, and on top of it, you can't force it for user as a condition for entering site (like Forbes is doing - they will get a complain, already beeing finalized by some privacy organisation)
Furthermore if the ePrivacy Regulation (ePR)[2], which was supposed to enter into force along-side the GDPR on May 2018 but was delayed, should be adopted in it's current form first party analytics like Matomo will not require consent. See [3]:
> The proposal also clarifies that no consent is needed for non-privacy intrusive cookies improving internet experience (e.g. to remember shopping cart history) or cookies used by a website to count the number of visitors.
[1] https://matomo.org/faq/general/faq_20000/
[2] https://en.wikipedia.org/wiki/EPrivacy_Regulation_(European_...
[3] https://ec.europa.eu/digital-single-market/en/proposal-epriv...
Also, who in their right mind would click "enable Google analytics" in opt-in mode?
Privacy opt ins are effectively donation buttons, and we know how well these work.
> Simpler rules on cookies: the cookie provision, which has resulted in an overload of consent requests for internet users, will be streamlined. The new rule will be more user-friendly as browser settings will provide for an easy way to accept or refuse tracking cookies and other identifiers.
[1] https://ec.europa.eu/digital-single-market/en/proposal-epriv...
> Opt-Out
I don't think thats how it works
How do you access data that's held about you by/in Google Analytics, anyway?
They just record on what you click, and have no (and will never have) idea of who you are.