K-anonymity
en.wikipedia.org
en.wikipedia.org
Follow-ups:
- k-map: https://desfontain.es/privacy/k-map.html
- l-diversity: https://desfontain.es/privacy/l-diversity.html
- δ-presence: https://desfontain.es/privacy/delta-presence.html
- differential privacy: https://desfontain.es/privacy/differential-privacy-awesomene...
https://github.com/KIProtect/data-privacy-for-data-scientist...
In the workshop we implement the "Mondrian algorithm" to produce a k-anonymous dataset. We then look at the problems of this approach (i.e. missing diversity in the sensitive attribute) and try to fix it using l-diversity (which is also not optimal) and finally t-closeness. The third notebook includes an implementation of a differentially private "randomized response" scheme, showing how it changes the data and how we can take into account the added noise when working with the randomized data.
I think it's important to keep in mind that k-anonymity and differential privacy are not algorithms but mathematical privacy definitions. To implement them, you need a suitable method like the "Mondrian" algorithm or a randomized response scheme.
If you have any questions or suggestions for improvements please open an issue or PR on Github!
https://www.troyhunt.com/ive-just-launched-pwned-passwords-v...
https://blog.cloudflare.com/validating-leaked-passwords-with...
https://www.okta.com/blog/2018/05/add-passprotect-to-your-we...
[1] Website: http://arx.deidentifier.org
[2] Source: https://github.com/arx-deidentifier/arx