No silver bullet: De-identification still doesn’t work
freedom-to-tinker.com
freedom-to-tinker.com
Very true. Furthermore, if you start to think seriously about making differential privacy guarantees across a whole company, it gets really hard and impractical.
It's not enough, as you might naively hope, to certify individual data analysis jobs separately as being "differentially private". Taken cumulatively they can amount to something which isn't [1]
It's not even sufficient to certify whole planned and delimited programmes of use for whole datasets as differentially private, if the company also holds (or wants to hold in future) other data on some of the same users and there's some chance of new actions being taken which directly or indirectly depend on both sources of data.
For the theory to truly apply, one needs to track (in a particular formal sense) how much "privacy budget" is used up by pretty much every user-data-dependent action taken across a company and across all datasets you hold pertaining to any overlapping set of users. Once your preallocated budget is used up it's pretty much game over in terms of taking any further actions based on any data (acquired now or in future) on any of those users, unless these actions can be based on inferences which were already obtained within the original budget.
These aren't necessarily insurmountable problems, and I'm sure there are new tools being developed to help which I'm not up to date with. (One which I am aware of is Microsoft's PINQ project [2]).
Still, it seems the organisations with the biggest chance of actually making this work, are those whose relationships with sets of users are inherently transitory and/or firewalled off from eachother. For example a B2B company who operate one-off surveys for clients, throwing away the raw data afterwards.
The fundamental problem is that differential privacy guarantees a degree of resistance to an incredibly strong adversary -- one with unlimited prior knowledge about your users, who is able to bring unlimited intelligence and computational resources to bear on drawing inferences about them based on every action you ever take conditional on user data. It's impressive that one can obtain any guarantees at all under this model, but perhaps not surprising that it can be hard to scale up and compose the guarantees without things blowing up.
That's not to say that there's isn't value in trying to obtain differential privacy guarantees for smaller scale pieces of work. It's still considerably better than other more naive pseudo-anonymisation methods.
[1] at least not to the same strength -- see http://en.wikipedia.org/wiki/Differential_privacy#Composabil... for the less handwavey version [2] http://research.microsoft.com/en-us/projects/PINQ/
What practitioners need is a canonical reference or toolset that you could feed a csv file to, tell it which columns should be scrambled and out comes an anonymized data set. Yes, I realize the "tell it" part is of concern, well that could be mitigated by smarter tools and more knowledgable practitioners - knowledge gained from a canonical reference.
We can't automate the process (in part because the transformations necessary are much more complex than "scrambling"), but knowledgeable practitioners can go a long way. I'm knowledgeable but not a practitioner, so I'm not the best source.
In #8 of our report (on the Heritage Health data), you'll notice that while I took Khaled El Emam to task for claims about quantifying risk, I do acknowledge that he did a very good job of de-identification. I don't think there's exactly a "canonical reference" (except HIPAA's superficial list of 18 identifiers), but reports written by practitioners like El Emam are probably useful documents.
Few people appreciate the robustness and generalizability of spatiotemporal coincidence analysis for reconstructing relationships in anonymized data both within and across many unrelated sources of anonymized data, even sources that are identifying entities that are not people (like anonymous vehicle tracking). There are enough anonymous entity tracking data sources available to algorithmically reconstruct relationships to most other "anonymous" data sources. I've seen it done many times and the capabilities are jaw-dropping in part because it violates human intuition as to what it is possible with such data sets.
Some of the notable results are:
* Golle and Partridge 2009 (http://xenon.stanford.edu/~pgolle/papers/commute.pdf) - Just the home and work location at city block level is enough to uniquely identify 50% of the US population. Home and work at zip code level is enough to uniquely identify 5% of US population. Home and work county is still enough to identify 1% of people to a set of 6 candidates.
* Montjoye et. al. 2013 (http://www.nature.com/srep/2013/130325/srep01376/full/srep01...) - If you have a time-location dataset of people with hourly accuracy of time and cell tower accuracy for location (100m in cities, a few km in rural areas), then four randomly picked points for each person uniquely identify 95% of them. Just two randomly picked points for each person uniquely identify 50% of people.
The main conclusion is that it is not possible to release an "anonymised" location dataset of people that is still useful for mobility research. The only realistic approach seems to be to have strict privacy regulations and NDAs with people who are given access to this data.
But more generally, all research into anonymisation and preventing de-anonymisation is difficult because it's not known what other data sources the attacker has access to. If I have a time-location dataset of anonymous people, and if I see from Facebook when a friend visited Paris and Barcelona, then it becomes trivial to match these dates and cities against the location traces, and find their full movement trace. Similarly, if I can get an anonymous dataset of phone calls, then I can make 20 missed calls at 4am to a friend, and later look for that pattern in the data to find their other calls.
The best example of this was from the Netflix dataset of anonymous movie ratings ( http://en.wikipedia.org/wiki/Differential_privacy#Netflix_Pr... ) - people were anonymous in their dataset, but some people also rated movies with visible identities on IMDB. Correlating the two datasets allowed researchers to discover the identities of people in the Netflix dataset.
I'm also generally skeptical about the level of damage caused by de-identification of intentionally released data versus the potential for damage from data released by hackers, subpoenas, and other means. Are there attempts to put these in perspective? Is a true adversary likely to be thwarted by better anonymization?
As a precursor to any honest conversation about an "equation", you would have to model the chance and impacts of re-identification. Anything else is sophistry to promote the idea of releasing data in general. Taking about how to anonymize data already carries the implicit assumption that there is value to be had in releasing it. Otherwise why would it be a topic of conversation at all?
I wouldn't have a problem with talking about the lack of quantification of this risk, it's a really hard problem if not impossible. What bothers me is that you're playing to a whole slew of fallacious reasoning, if I was being less charitable I'd say you're trying to go congress on "the debate."
My prejudice is certainly toward the release of data in the case of datasets of clear use to research that lack obvious harm to the participants. I feel that because of randomwalker's work, datasets such as those used for the Netflix Prize may never be released again, and I feel that this is a net loss to society: https://news.ycombinator.com/item?id=1193417
What bothers me is that you're playing to a whole slew of fallacious reasoning, if I was being less charitable I'd say you're trying to go congress on "the debate."
I'm not familiar with the phrase "go congress". Could you explain? I'm not trying to play to fallacious reasoning, although I certainly might be susceptible to it. I genuinely would like more discussion of what is lost when data is not released due to fears of violation of privacy. Could you be more specific about the fallacies involved?
Is this just a random feeling or is there something to back this up?
How are you even assessing things like the value of the difference in precision caused by doing something like creating (mostly) distribution equivalent (but fake) data and the cost of leaked person details to hundreds of thousands or millions of people?
The case I'm most familiar with is the Netflix Prize. I think I can safely say (in my semi-professional opinion) that a lot of good research was published as a result of Netflix's decision to release the data that they did: http://scholar.google.com/scholar?q=netflix
I view these publications as a public good. It's possible the techniques described were previously known privately, but before the contest there was no description of them in the available literature. At the conclusion of the contest, Netflix briefly announced that they would have a second contest, which would involve the release of another data set. In large part as a result of the press coverage of randomwalker's work, this potential release was cancelled: https://freedom-to-tinker.com/blog/paul/netflix-cancels-netf...
It's disputable, but I feel this additional data set would have generated a similar number of good publications, inspired new research, spread knowledge, and offered a similar benefit to the public.
It can be argued that preventing further data releases has prevented further harm to the public. But I think it's important to look at the actual harm done by the release of the first dataset. While there was some degree of potential for harm, to my knowledge no individual was actually discriminated against as a result of being identified in the data set. By contrast, I think it is accepted that numerous individuals have been harmed by other online data breaches. If more attention is paid to anonymous datasets than to other matters that cause actual harm, the attention is likely misguided.
I'd draw a parallel with the increase in airline security post-9/11. It's possible that confiscating oversize toiletry items and removing shoes has prevented further harm. It's indisputable (I think) that the additional hassle has had a negative cost to the public. And if more people choose to drive than to fly, the net effect is likely negative. But at least in the case of airline security, there is a clear case to be made of actual harm. Whether the current policies are a good compromise of inconvenience and safety depends on both the risk and the reward.
I know less about the other cases, both for benefit and harm. Have New York taxi drivers suffered harm as a result of the poor attempts at anonymizing the data? Is there public benefit to the data that was released? Have individuals been harmed by the semi-anonymized release of health records? Was this offset by any positive effects? This is the discussion I think we should be having, rather than simply pointing at the abstract potential for harm.