'Anonymised' data can never be totally anonymous, says study
theguardian.com
theguardian.com
Also related: https://news.ycombinator.com/item?id=20513453.
To clarify, when de-identifying a dataset you simply remove direct identifiers (like a name) from it. This protects the data from direct re-identification, i.e. from someone learning the identity of a person in the data by looking at an individual row. Anonymization is supposed to protect individuals from re-identification also when using external context information and can usually only be achieved by further transforming the data, for example by grouping it (as techniques like k-anonymity do), adding noise (as randomized response techniques do) or by synthesizing new data.
High-dimensional, de-identified data will always be easy to re-identify given enough context, I've done this myself in 2016 with a clickstream dataset (the authors reference our work in their paper).
https://medium.com/capital-one-tech/why-you-dont-necessarily...
It’s even good enough in many cases to build models off of.
However, it’s important to not that numbers such as SSNs are always “real”, in the sense they link to someone. A random SSN has with near certainty been used before and belonged to someone (perhaps multiple) people. The trick is ensuring the rest of the attributes don’t match the individual.
Tangent: I suppose humans choice in models and statistics to use can be argued to be a similar thing, but I'm still dubious of how this concept is being communicated.
What we do is extract features in an unsupervised way and build a model to recreate the dataset from said feature. When the synthesized dataset is indistinguishable from the real thing. You should be able to build a model that gets you 90%-100% of the way there with synthetic data, then tune / retrain on the real data.
The model that does decisioning itself doesn’t matter. If you needed to adjust the synthetic dataset, you’d have to either build a new generative model OR otherwise weight the dataset.
Another issue I have with data synthetization is that by trying to reproduce the original data as faithfully as possible you waste a lot of your privacy budget on meaningless features.
That said I'm happy to be convinced otherwise. Do you have a publicly available demo dataset that I could look at?
Unmasking of bulk-sold pseudonymized user data is an externality, like pollution — those who bear the cost when the data gets reidentified are the users, not the buyers or sellers. Therefore the data belongs to the users and propagation should be severely constrained.
Anonymity is more polite fiction than rigorous fact these days.
Let me give a concrete example. Let's say you want to collect application usage information from users of your app. You could a) collect all the information attached to an IP address, and any other nominal information, and store it, or b) you could calculate the valuable usage metrics for the entry, and store it in some demographic bucket, throwing away any other knowledge.
The real problem is that companies want the full data so they can be free to change their models at will. We as users should not expect data to be stored properly, and cryptography tools are our main tool to address this problem.
A couple years ago, I got into an argument on reddit where someone claimed that any mapping could be recovered "using deep learning techniques" (e.g. if you take 3*0 = 0, you can get back that the original value was 3 with no other information except for the value "0"), and that obviously I was just too stupid to understand deep learning if I couldn't see that.
On a related note, while x * 0 effectively erases knowledge of x, something like x ⊕ s, where s is some secret, reversibly obscures it.
The notion of reversibility is the key misunderstanding.
And it's not intuitively obvious which combinations of values allow you to recover which other ones.
> And it's not intuitively obvious which combinations of values allow you to recover which other ones.
I think it's pretty intuitive that Zip Code and DOB are identifiers. That's why they count as such in HIPAA, and are used to demonstrate identity by governments, credit cards, etc.
Personally I think this stuff just poisons the well when it comes to discussions of privacy. I think the goal is to remove the expectation of anonymity by claiming that it's never possible.
It's great that you think that, but basically no company uses that definition. Most company privacy policies don't consider combinations of information when making this determination. E.g. your billing address might be personal information, but your zip code by itself might not. Similarly, IP address (with or without last octet), wifi SSID, location data, browsing history (or attributes derived from browsing history), and so on. Each individual piece of data isn't enough to personally identify you, so the privacy policy often doesn't have to be applied to it.
E.g. after reading the Google privacy policy[0], can you tell what protections your zip code and DOB have? Will Google treat them as personal information or personal identifiers or not?
Sure, but what about job title? What about job title when someone's job title is "mayor" or "fire chief" and it's possible to deduce from other information what city they're in? Or someone's job title is just "Governor of the State of California"?
Any collection of random independent characteristics become uniquely identifying once you have enough of them. Then all the attackers need is another database with the same characteristics that also includes names or other identifiers, and you can associate the missing fields from one database with the other.
> Personally I think this stuff just poisons the well when it comes to discussions of privacy. I think the goal is to remove the expectation of anonymity by claiming that it's never possible.
It's not that it's never possible, it's that it's only possible if we don't feed these centralized databases enough information to uniquely identify people. So we need to stop doing that.