Differential privacy tools from MS Research and Harvard
blogs.microsoft.com
blogs.microsoft.com
"Differential privacy is the gold standard definition of privacy protection."
Differential privacy is nice, but it's still tracking and "less tracking" is absolutely not the gold standard definition of privacy protection. That title goes to "no tracking"
Things like emoji usage, page navigations, feature uses, etc. Ideally anonymous; no IP, user agent, etc, just a small byte or two packed properly can go a long way.
The problem is that it's actually quite hard to reliably anonymize data especially once you start to begin combining data sets from multiple places. That's the problem differential privacy is trying to solve in a mathematically rigorous way.
See for example how researchers partially de-anonymized Netflix Prize data by cross-referencing it with IMDB reviews.
DDG:
{ used_advanced_search }
{ used_country_toggle }
{ tabbed, *tab_maps }
{ filtered, *filter_date }
{ os "iOS", *ver "13.5", browser "Safari" }
iOS: Mail
{ disabled_remote_images }
{ flagged_mail }
Keyboard
{ emoji_keyboard_via_globe }
{ *emoji_use "100-1000", *emojis [ ":)" ":P" ":(" ] }
Each of these could be stored separately without metadata then aggregated no problem. Things marked * could be left out, and some things could be randomized up or down buckets and such.Right now you're probably wondering, yeah, but there's this one problem that wouldn't have been solved if x or y....
But really it's just hoarding behavior. They're trying to collect it all. Statistical significance is reached very quickly and after that point they're doing harm to society.
"Everything you say can and will be used against you"
All they study is ways you are bad, and all they research is ways to keep you down.
If anyone reaches the wrong conclusion "maybe they are innocent.." Then it only takes two seconds, then they are fired.
Privacy from social scientists is one of the most important forms of privacy.
This is roughly analogous to "abstinence is the best form of birth control". It's not wrong, but it also isn't particularly realistic or helpful. The reality is that people often do want to exchange data, some times because it is legally or morally mandated, and tools that allow this to be done as safely as possible are important.
If you have free cycles, you can read more here:
https://github.com/frankmcsherry/blog/blob/master/posts/2017...
This depends on how many researchers and queries per researcher that you want to allow. The privacy budget eventually runs out, so there are definitely drawbacks that prevent effective use of this data across enough outside researchers.
this is why DP doesn't get used in any real system -- limited # of searches is a deal breaker for any service that wants to monetize
there are some applications where this could be okay, like in-company surveys where you want to enable employees to run stats queries without revealing individuals, but companies are (relatively) high trust environments and DP is overkill
It was used for the 2020 US Census.
There are some techniques like exposing randomized subsets for a limited number of queries.
how is access to the privacy prioritized?
Also, there are implementations of DP that do not rely on limiting queries. Lastly, most companies aren't "high trust" environments. There are tons of stories about employees routinely abuse their company's data. Companies with any sensitive data should be looking at DP and other anonymization techniques even if their data is only ever shown to their employees.