Too Unique to Hide
cpg.doc.ic.ac.uk
cpg.doc.ic.ac.uk
When I was in the UK I had to travel to a B&B at night. I had a GPS but they hadn't given us a street address - just a postcode and house name. You couldn't enter that into our ancient, so I rang them. Initially they didn't give me an address, telling me it was useless. When I persisted they gave me a street name and suburb. I forget the suburb, but I will never forget the street name. It was "The Street".
I entered what I had in the GPS, and two hours later I arrived where it said to go. But there was no house with the name given. So we rang them again. Turns out there were three streets named "The Street" all within 1km of each other and we were in the wrong one.
We arrive at the correct "The Street", but still can't find the house. We ring them for help, but it's freezing and raining, or at least doing what passes for rain outside of London and they are reluctant. Admittedly it's a short street. We feel like fools, but as we still can't find it so keep calling, and eventually our hosts stand on the road in the rain and flag us down. Turns out the B&B is on a battle axe block and the car has to push past bushes in the entrance to get onto the driveway.
At breakfast I make the mistake of expressing my opinion of the UK's street address system. They look insulted and defend it vigorously. Then as we load the car I hear them having a loud "discussion". It turns out a new postman has been appointed who hasn't learnt the house names yet, so the mail hasn't arrived for a few days. They were discussing who should go to the post office to collect it.
https://www.hitta.se/kartan!~57.68616,11.91682,16z/tr!i=5N3q...
I understand it's for a reason (see here https://hejsweden.com/en/swedish-privacy-personal-informatio...) and I respect it, but I wonder how long it will take until someone makes bad usage of these data.
Here's an example: GL52 2SE
The GL52 part covers large parts of Cheltenham.
https://www.postcodes-uk.com/GL52-postcode-district
But the full post code GL52 2SE covers only 8 addresses in Albion Street. The full post code covers from KwikFit up to 132.
https://www.google.co.uk/maps/place/Albion+St,+Cheltenham+GL...
Plus this bike shop: https://www.google.co.uk/maps/place/Albion+St,+Cheltenham+GL...
If only people would practice what they preach.
t-closeness, is defined as: An equivalence class is said to have t-closeness if the distance between the distribution of a sensitive attribute in this class and the distribution of the attribute in the whole table is no more than a threshold. A table is said to have t-closeness if all equivalence classes have t-closeness.
In short: the distribution of a particular sensitive value should not be further away than a distance t from the overall distribution.
Using the t-closeness metric circumvents issues associated with k-anonymity and ℓ-diversity. Briefly, k-anonymity states that a certain attribute class should be present in at least k records, which introduces ambiguity in the data set. However, if each of the k equivalence classes are the same, properties could still be resolved simply by elimination. The ℓ-diversity metric circumvents this problem by adding a further requirement: in addition to the class to being seen in k records, these records must have at least ℓ ‘well represented’ values. But if an attacker knows the real-world distribution of values, then attributes could still be disclosed with a certain probability, simply by combining different data sources
Yes it bloody is...
Continuing anyway, I arrive at "NaN% of the time!"
We used the complete list of US/UK postcodes (edit: England and Wales only for UK) to implement the demo, so all of them should be there.
1) Consider that your date of birth + postcode appears twice in a dataset. Naively, the operator can think to correctly identify you 50% of the time (you are either entry #1 or entry #2). But, it is also possible that you are not actually in the list. Maybe there are two other people in the area with your DoB which are not in the database, thereby reducing the odds of guessing correctly. You have to correct for this sampling bias. Also, the result is now not necessarily in the form of 1/n anymore.
2) Consider that you now have two datasets that together contain the whole population (but there is overlap between the datasets that is difficult / impossible to remove). Your DoB + postcode appears 3 times in the first set and 5 times in the second dataset. Then ask again: what are the odds of identifying you from the 8 candidates? Again, this does not give a solution in the form of 1/n.
There was one poor individual who was the only person who lives in a certain post code area.
Ended up having to get the dataset owners to remove postcodes altogether (GDPR just came in and everyone was on high alert).
It's funny how in my head, it seems like it would be way less likeley that you could be indentified with such a high degree of confidence just from those 3 peices of info.
[0] - https://en.wikipedia.org/wiki/BT_postcode_area#Coverage
And for Scotland from NR Scotland - https://www.nrscotland.gov.uk/statistics-and-data
Edit: changed.
E.g. If there’s 1 female in post/zip code XXXX born in 09/79, doesn’t matter if you’re only including the month. There’s a good chance you can work out who she is.
It’s a sliding scale dependent on the other attributes in the dataset, and the number of distinct values across all the attributes.
The only way to make data completely anonymous is to remove things like birth date tbh.
If you can't restrict to a specific geographic area, then who knows who you are.
Pragmatically speaking, I noticed that in their "test toggling variables" widget at the end, removing postal code instantly dropped my probability of detection to 0, no matter what other fields were active.
So perhaps your geographic location is the information you should be most paranoid about sharing.
> Where do you live?
> UK US