Is anyone aware of any work for how to apply differential privacy to language models?
So the main question I have is let's say I'm working with sensitive data like emails or doctors notes. How can I train an ML model that would still learn something useful without leaking private data.
When I say "leak", an example would be I train an RNN on some company data email data and when I feed the RNN "$AMZN" the network would say SELL.
How can I quantify how much the model has learnt and how much privacy has been leaked.