What practitioners need is a canonical reference or toolset that you could feed a csv file to, tell it which columns should be scrambled and out comes an anonymized data set. Yes, I realize the "tell it" part is of concern, well that could be mitigated by smarter tools and more knowledgable practitioners - knowledge gained from a canonical reference.
We can't automate the process (in part because the transformations necessary are much more complex than "scrambling"), but knowledgeable practitioners can go a long way. I'm knowledgeable but not a practitioner, so I'm not the best source.
In #8 of our report (on the Heritage Health data), you'll notice that while I took Khaled El Emam to task for claims about quantifying risk, I do acknowledge that he did a very good job of de-identification. I don't think there's exactly a "canonical reference" (except HIPAA's superficial list of 18 identifiers), but reports written by practitioners like El Emam are probably useful documents.
Few people appreciate the robustness and generalizability of spatiotemporal coincidence analysis for reconstructing relationships in anonymized data both within and across many unrelated sources of anonymized data, even sources that are identifying entities that are not people (like anonymous vehicle tracking). There are enough anonymous entity tracking data sources available to algorithmically reconstruct relationships to most other "anonymous" data sources. I've seen it done many times and the capabilities are jaw-dropping in part because it violates human intuition as to what it is possible with such data sets.
Some of the notable results are:
* Golle and Partridge 2009 (http://xenon.stanford.edu/~pgolle/papers/commute.pdf) - Just the home and work location at city block level is enough to uniquely identify 50% of the US population. Home and work at zip code level is enough to uniquely identify 5% of US population. Home and work county is still enough to identify 1% of people to a set of 6 candidates.
* Montjoye et. al. 2013 (http://www.nature.com/srep/2013/130325/srep01376/full/srep01...) - If you have a time-location dataset of people with hourly accuracy of time and cell tower accuracy for location (100m in cities, a few km in rural areas), then four randomly picked points for each person uniquely identify 95% of them. Just two randomly picked points for each person uniquely identify 50% of people.
The main conclusion is that it is not possible to release an "anonymised" location dataset of people that is still useful for mobility research. The only realistic approach seems to be to have strict privacy regulations and NDAs with people who are given access to this data.
But more generally, all research into anonymisation and preventing de-anonymisation is difficult because it's not known what other data sources the attacker has access to. If I have a time-location dataset of anonymous people, and if I see from Facebook when a friend visited Paris and Barcelona, then it becomes trivial to match these dates and cities against the location traces, and find their full movement trace. Similarly, if I can get an anonymous dataset of phone calls, then I can make 20 missed calls at 4am to a friend, and later look for that pattern in the data to find their other calls.
The best example of this was from the Netflix dataset of anonymous movie ratings ( http://en.wikipedia.org/wiki/Differential_privacy#Netflix_Pr... ) - people were anonymous in their dataset, but some people also rated movies with visible identities on IMDB. Correlating the two datasets allowed researchers to discover the identities of people in the Netflix dataset.
Very true. Furthermore, if you start to think seriously about making differential privacy guarantees across a whole company, it gets really hard and impractical.
It's not enough, as you might naively hope, to certify individual data analysis jobs separately as being "differentially private". Taken cumulatively they can amount to something which isn't [1]
It's not even sufficient to certify whole planned and delimited programmes of use for whole datasets as differentially private, if the company also holds (or wants to hold in future) other data on some of the same users and there's some chance of new actions being taken which directly or indirectly depend on both sources of data.
For the theory to truly apply, one needs to track (in a particular formal sense) how much "privacy budget" is used up by pretty much every user-data-dependent action taken across a company and across all datasets you hold pertaining to any overlapping set of users. Once your preallocated budget is used up it's pretty much game over in terms of taking any further actions based on any data (acquired now or in future) on any of those users, unless these actions can be based on inferences which were already obtained within the original budget.
These aren't necessarily insurmountable problems, and I'm sure there are new tools being developed to help which I'm not up to date with. (One which I am aware of is Microsoft's PINQ project [2]).
Still, it seems the organisations with the biggest chance of actually making this work, are those whose relationships with sets of users are inherently transitory and/or firewalled off from eachother. For example a B2B company who operate one-off surveys for clients, throwing away the raw data afterwards.
The fundamental problem is that differential privacy guarantees a degree of resistance to an incredibly strong adversary -- one with unlimited prior knowledge about your users, who is able to bring unlimited intelligence and computational resources to bear on drawing inferences about them based on every action you ever take conditional on user data. It's impressive that one can obtain any guarantees at all under this model, but perhaps not surprising that it can be hard to scale up and compose the guarantees without things blowing up.
That's not to say that there's isn't value in trying to obtain differential privacy guarantees for smaller scale pieces of work. It's still considerably better than other more naive pseudo-anonymisation methods.
[1] at least not to the same strength -- see http://en.wikipedia.org/wiki/Differential_privacy#Composabil... for the less handwavey version [2] http://research.microsoft.com/en-us/projects/PINQ/