Court confirms that IP addresses are personal data only in some cases (2016)
whitecase.com
whitecase.com
You also, by the way, must remove any geographic information more specific than a state, such as ZIP codes. So it doesn’t say much to include IP addresses in the deidentification list.
That's the detailed version of the ruling. The ruling refers to IP adresses with time and date, as explained in point 37.
So IPs and time stamps are not PII by themselves.
>a dynamic IP address registered by an online media services provider [...] constitutes personal data within the meaning of that provision, in relation to that provider, where the latter has the legal means which enable it to identify the data subject with additional data which the internet service provider has about that person.
I think such reasoning is a little unfortunate.
Multiple people share my name. Multiple people could live in or visit my household.
So is my name and address not PII?
I don't think that underlining all the cases where some given bit of information may fail to identify a person is the right approach when it comes to making a blanket statement about whether said info is PII. I don't think courts would follow that reasoning either, especially when there will be lots and lots of counterexamples where following that trail of information leads to facts that most people would conclude as identifying a person exactly (at least with a very high degree of certainty).
> Examples of personal data
> [...]
> an Internet Protocol (IP) address;
Source: https://ec.europa.eu/info/law/law-topic/data-protection/refo...
So the title on HN is misleading. Especially given the most important part of the article:
"The CJEU decided that a dynamic IP address will be personal data in the hands of a website operator if:
- there is another party (such as an ISP) that can link the dynamic IP address to the identity of an individual; and
- the website operator has a "legal means" of obtaining access to the information held by the ISP in order to identify the individual."
And it's known that the "legal means of obtaining access to that information" is very often present.
So an IP address on its own is almost never personal data, because of wifi, NAT, dynamic IPs, shared devices, etc. Then again, a name is almost never personal data on its own either, "John Smith" could refer to any one hundreds of thousands or people or it could be a pseudonym and refer to literally billions of people.
But if someone registers on your site, and you log the IP address and their name, you're a lot closer to persona data. Add a timestamp, and you probably can identify a real person.
So if you're trying to be careful about GDPR, you should probably be careful about storing IP addresses (or IP addresses that can be linked to other bits of potentially personal data). The focus of GDPR compliance can't be on "oh this field is fine, but this field is personal data", it should be on what you're collecting in aggregate. That makes IP addresses dangerous, because they provide a lot of information that could be used to identify someone.
So based on my reading, IPs and time stamps are not PII unless you are an ISP or you link them to other PII (so still the IP and time stamps are really irrelevant because they depend on that other PII).
A web access-log records (ts, ip, request, ...), or maybe your application log stores (ts, ip, action, params, ...)
So the information from that single source is "at time T, IP accessed RESOURCE".
It's possible that's personally identifiable in context (if you have additional controls that RESOURCE can only be accessed by exactly 1 real person, etc)
But say it's not. All you know is: Opaque PERSON accessed RESOURCE.
if you can obtain the identifying information from elsewhere (buy, steal, etc) from ISP or whatever, you now know that (T, IP) = NAMEDPERSON.
A simple lookup/matching means you know that NAMEDPERSON accessed RESOURCE. That's the new personal data.
The IP isn't irrelevant, because without it, you'd have no lookup key to determine the mapping from PERSON? to NAMEDPERSON.
If an ISP is willing to sell that data, are IP addresses now PII for everyone?
If one part of a company has such a DB, does it apply to every part of the company? What if it's multiple companies owned by a conglomerate?
If you include an image (or a font!) from somewhere else in a web page, you are causing the user's IP address to be sent to the hosting party, are you liable for sending PII if the target can link IPs to names, because they (e.g. Google) have a DB?
As an end user, I want this—if you wouldn't send my IP address to these people otherwise, wanting to show me an image or a font is not a good reason to send it.
As a web developer, I am happy to have excuses to tell my teammates that we need to rehost every asset we depend upon. It's the right thing to do for so many other reasons.
I know this makes things hard for people who have webfonts that don't allow rehosting them etc. Being able to say "We can't use this font because of GDPR unless you change your policies" sounds pretty great honestly.
On the topic of proprietary fonts, why do some websites seem to think using a questionably legible, licensed font is a good idea? It isn't adding value for the end user.
Consider: You're CompanyX, and I'm DodgyFontHost.tld My business model is exploiting and selling as much data as I can gather/mine from my traffic.
You embed (that is, reference/hotlink) some of my fonts on your pages.
If a user visits your page, and as a result makes a request to me for a font, I can log everything about that request, but I don't (afaik) have much/any additional knowledge that makes it particularly useful.
Assume there's no ?UTM=... tracking content in the url itself, you're just referencing a static font file.
I'm not sure offhand if browsers would be passing a referer header by default, or if that could somehow reliably identify the site I'm actually visiting. If so, that'd be one valuable fact.
I might be able to fingerprint the users browser from other headers or their OS from network-level quirks.
Anything else I'm missing?
I feel like 'IP $x made a request for $file' isn't the important thing to be looking at here, it's what I can learn from other things associated with the request that I can exploit.
But yes, if you had a reliable lookup from (ip,timestamp) to legal person, then it's absolutely Personally Identifiable.
Imagine if every browser set a valid, correct 'X-Requestors-Legal-Name: Bob Smith, Sometown, USA' header on every request. That's obviously identifiable. Adding a layer of indirection doesn't make it less so, although it does maybe place it on a continuum of 'cost/effort to identify based on this info'.
It ranges from 'trivial, because it's right there in the content you're sending', through 'not directly, but easily enough via subscriptions to one or more commercial data providers' to 'if someone steals our data and combines it with stolen data from several other sources, they have a non-zero chance of guessing your identity correctly'.
They are passing referer unless the context is an encrypted connection and the resource is on plain HTTP.
Firefox also strips out the path from the URL for third party requests, but only in private browsing mode: https://blog.mozilla.org/security/2018/01/31/preventing-data...
I think this should be the default for all third party domains no matter what the mode. (Really, I'd rather see that header just go away.)
I suspect it's sufficiently ingrained in existing apps to make it hard to deprecate completely, but something like the path stripping might be a decent compromise.
For cross-origin requests I think there's also a mandatory 'Origin:' header that would identify at least the domain (but not path) a user request was referenced from.
I used to use a firefox addon called RefControl but IIRC it was a casualty of the quantum/webextensions transition. uMatrix has a basic referer spoofing capability, but it's all or nothing for a particular site/scope.
I agree with others that you should self-host whenever possible. It will simplify these questions and you'll be able to fully protect your users' data yourself.
> However, businesses should note that if they have sufficient information to link an IP address to a particular individual (e.g., through login details, cookies, or any other information or technology) then that IP address is personal data, and is subject to the full protections of EU data protection law.
Do a geolookup, you have my approximate location.
Do a Google search for my IP address and you'll have my name.
IPs specifically are quite likely to reveal some identifying info, and it's obvious how trivial it is to find that info. Even the company itself isn't looking that info up, losing that info could expose their users.
How does that happen exactly ?
Name is required on the whois record, but even if it could be anonymized, it'd still have the registrar's name. I am my registrar.
From my reading, you can interpret it that way if "the website operator has a 'legal means' of obtaining access to the information".
Refer to the "What makes a dynamic IP address personal data?" section of TFA.
(N.B.: I am not a lawyer. Ask your doctor if taking legal advice from strangers on the Internet is right for you.)
Nowhere near "personally identifiable" nor necessarily correct in any way.
> Do a Google search for my IP address and you'll have my name.
That would be highly unusual.
I'm not sure about the other RIRs but ARIN, at least, has (had?) a requirement that any assignment of a /29 or larger must be reported (see "SWIP" [0]).
In other cases, a PTR RR for a single IP address could be enough to personally identify an individual.
Commercial entities doesn't really map to one person. I thought WHOIS would have to be amended to be compatible with GDPR anyway.
I mean, you could create a website dedicated to mapping your current IP to yourself if you really wanted to, but that is hardly relevant.