Is Writing Style Sufficient to Deanonymize Material Posted Online?
33bits.org
33bits.org
[1] http://news.ycombinator.com/item?id=413730
[2] No, really, it's been exactly Π years to the day :-)
As a specific example, people writing political blogs in China could be seriously harmed by this technique even at the levels that it's at now.
I applaud you for including the link to "manually changing your writing style will defeat these attacks" but that's a link to an academic paper. Could you please also write some good, layperson-oriented docs on "how to beat this"? For that matter, I'll do the writing grunt work if you'll provide the expertise. If you're interested, use the GMail address in my profile.
As for practical tips to defeat stylometry and such, organizations like the EFF specialize in doing that, so I will leave that to them. Comparative advantage, etc. If you would like to help, you are more than welcome.
There's an interesting presentation here: http://events.ccc.de/congress/2011/Fahrplan/attachments/2019...
I'd be interested in seeing more tools in this vein. I already know that if I want to camouflage my writing, I have to cut way down on the dashes, but it'd be neat to upload a few large documents and see a nice list of the top ten traits that could be used to identify me.
Making this technology available and easily accessible for widespread public use encourages development of counter-authorship-detection techniques. If tools are available that allow authors to easily assume the identity of other people (through morphing of writing styles), this technology rapidly loses value.
Ask HN's are awesome but you don't usually get a PhD thesis in reply :-)
Imagine I want to post anonymously about some sensitive subject. Whenever I do it, I write the article normally, then Google translate from English to French to English, and then I clean up obvious errors in the retranslation.
Theoretically it's not so easy. If you know how Google translates from English to French (and vice-versa) you can at least partially reverse the process.
Or perhaps authorship detection can be performed on the output from Google translate? The authoriship entropy contained within the original English text is likely to be carried through (at least partially) to the post-translated output.
If so, perhaps it would be possible to use it to engineer countermeasures. In the example you used ("since" vs "because"), it would be fairly simple to alter the ratio.
Of course I don't know what more complex indicators you may be using, but I'm having a hard time imagining what would be easy to measure, but difficult to alter.
I guess you did not look at texts with multiple authors, or professionally edited texts? I am curious if a different editor, or publication house style, can be detected.
Before the publication of the manifesto, Theodore Kaczynski's brother, David Kaczynski, was encouraged by his wife Linda to follow up on suspicions that Ted was the Unabomber. David Kaczynski was at first dismissive, but progressively began to take the likelihood more seriously after reading the manifesto a week after it was published in September 1995. David Kaczynski browsed through old family papers and found letters dating back to the 1970s written by Ted and sent to newspapers protesting the abuses of technology and which contained phrasing similar to what was found in the Unabomber Manifesto
It reminds me of the claims of being able to identify, for example, the gender of an author with ~65% accuracy -- which is really actually completely unimpressive, as it's hardly better than guessing, and certainly not something you could rely on for any serious purpose.
The author mentions that topic is one way to help correlate beyond the results of the algorithm. But if I wrote "anonymous" posts in my area of expertise, you certainly would not need stylistic analysis to guess what my identity might be! There has never been privacy in this regard, I don't think.
Where privacy is needed most, I think, is exactly where this deanonymizing tool still isn't sufficient: talking about unrelated topics. A person should be free to express themselves under multiple names for different purposes, and there is no reason why an employer needs to know about a programmer's side hobby as a fiction writer if s/he doesn't want them to.
Finally, I do wonder how well these results correlate to the case where someone is intentionally operating under a different name. Matching one post by tech blogger A against blogger A is easy, because tech blogger A is making no attempt to write any differently or in any different context. However, what if tech-writer A ghost-wrote YA fiction on the side? Could you use these techniques to detect that the fiction was written by that blogger? It can't be ruled out without trying, but generalizing these results to that seems questionable.
One thing that'd be interesting to me is whether there are certain characteristics that make it particularly easy to identify people cross-context, like a top-10-telltale-markers sort of thing. Are a disproportionate number of the 10% who can be identified with high precision using a handful of unusual grammatical or lexical features, or is it more of a diffuse sort of thing?
http://ata-s12.utcompling.com/schedule/ATA-Authorship%20Attr...
It seems you didn't attempt to fingerprint for misspellings, among the variables on pdf p 5. Also, curious why did you need to up the dataset to exactly 100k with the 5.7k.
http://www.cmu.edu/news/archive/2011/January/jan7_twitterdia...
Maybe running your text through a round-trip translator could help? (although then you'd need to fix any errors introduced).
If your writing is too identifying, just perturb the text until the tool fails to identify the author. Or even better: perturb the writing until the deanonymizer fingers someone else, in a usefully confounding way.
The deanonymizer's feature-extraction/analysis could itself help drive the perturbation routines. "Make my word choice more like Paul Graham", you could say. And even if there are limits to its automatic substitutions, it could offer coaching: "To make your writing more Graham-like, decrease your average sentence length and use fewer interjections."
Edit, resubmit, repeat until the right author is fingered.
Business idea: website that offers this tuning to help un-deanonymize or faux-deanonymize writing.
Evil business idea: this website remembers everything submitted, to allow the super-deep-pocketed to peek in and de-un-deanonymize (of re-deanonymize?) blocks of text.
Method1: Run the text through a markov chain constructed maybe from a mixture of 0.5 your text, 0.25 Shakespeare and 0.25 Alice in wonderland. Do something like sample every third word with the other two coming as a chain. Then run that text through wordnet to do synonym based replacement.
Method 2: Do a translation to a nearby language and back again using some language translating api.
Method 3: Replace less common words with hypernyms and more common words with synonyms or possibly not + antonyms.
Might want a few heuristics to replace stuff like (, ..., ) , - , : ,[,] with each other. Also randomize space between punctuation.
Optionally Run the outputs through mechanical turk to iron out the result, leave as is or clean by self.
Of course, this is dependant on translation algorithms that are at least somewhat inaccurate.
You may want to choose one closeby language, and one further removed and find the equilibrium phrase. For example, translate english->german->italian->german->english, repeat until you get the same english phrase each time.
I am consistently astounded by how advanced AI techniques are becoming.