358 karma · joined April 5, 2017
But most probably I'll be publishing a paper later this year detailing all the details and process :)
So as basis for my thesis on AI and NLP I've been working on a RRN-based text classifier that basically reads and analyzes privacy policies. It understands that "we don't share your data with third parties" is privacy friendly while "we may share your data with anyone" is a potential threat.
I've then created this website with a bunch of analyzed services to showcase the most relevant info about each service along with other interesting stuff like recent data breaches or instructions to delete your account in said service.
Happy to answer Qs about the tech behind, it'd also be great to hear your feedback on what the site lacks and possible improvements!
I’m working on an AI that reads privacy policies and automatically detects privacy threats for you.
Last year I started becoming really concerned about digital privacy when I found out Facebook has always had an updated copy of my phone contacts, including nicknames and notes [1] – which basically means complete strangers now know the names I call my gf and the notes I would put on people to remember them (ex: John – that creepy guy from work) because I stored them like that. It was a total invasion of my privacy, it was too much.
But it was not so surprising, because later on I found out that Facebook explicitly says they’ll do this in their privacy policy. That same privacy policy no one reads.
So when I had to choose the topic to write my CS thesis, I was pretty sure I wanted to choose privacy – and as an AI enthusiast, I set myself to solve this problem by teaching machines to defend us, so we could all enjoy a safer internet.
After some research and training data, I managed to create a Recurrent Neural Network that tells apart potential privacy threats in policies (like “we will sell your data”) from neutral or privacy-friendly sentences (like “we anonymize your data even before it reaches our servers”).
I’m building this for people that value their privacy. I think you guys might appreciate it, so I’d love to hear what you think. The main problem now is that, as in any deep learning model, it needs a huge amount of data to be trained accurately. The AI is not very precise right now. You can help push privacy as a larger concept forward just by playing a simple game that (a) will let you know some of the biases you may have when you face privacy “dilemmas” while (b) it will further teach the AI what’s privacy friendly and what’s not [2].
Right now this is an academic experiment turned webapp (we’ll probably publish a paper later this year) – but in the future I’d like to build a set of tools that actively act as countermeasures to these privacy threats [3], and I’d like to monetize the project that way.
[1] https://news.ycombinator.com/item?id=16661735 [2] https://privacyreader.useguard.com/experiment [3] https://useguard.com
As far as I know, living in the EU does help: my understanding is that within the EU, all your data belongs to you even if it's stored by Facebook; as opposed as the US, where Facebook is the rightful owner of the data you provide.