[1] http://news.ycombinator.com/item?id=413730
[2] No, really, it's been exactly Π years to the day :-)
[1] http://news.ycombinator.com/item?id=413730
[2] No, really, it's been exactly Π years to the day :-)
As a specific example, people writing political blogs in China could be seriously harmed by this technique even at the levels that it's at now.
I applaud you for including the link to "manually changing your writing style will defeat these attacks" but that's a link to an academic paper. Could you please also write some good, layperson-oriented docs on "how to beat this"? For that matter, I'll do the writing grunt work if you'll provide the expertise. If you're interested, use the GMail address in my profile.
As for practical tips to defeat stylometry and such, organizations like the EFF specialize in doing that, so I will leave that to them. Comparative advantage, etc. If you would like to help, you are more than welcome.
There's an interesting presentation here: http://events.ccc.de/congress/2011/Fahrplan/attachments/2019...
I'd be interested in seeing more tools in this vein. I already know that if I want to camouflage my writing, I have to cut way down on the dashes, but it'd be neat to upload a few large documents and see a nice list of the top ten traits that could be used to identify me.
Making this technology available and easily accessible for widespread public use encourages development of counter-authorship-detection techniques. If tools are available that allow authors to easily assume the identity of other people (through morphing of writing styles), this technology rapidly loses value.
Imagine I want to post anonymously about some sensitive subject. Whenever I do it, I write the article normally, then Google translate from English to French to English, and then I clean up obvious errors in the retranslation.
Theoretically it's not so easy. If you know how Google translates from English to French (and vice-versa) you can at least partially reverse the process.
Or perhaps authorship detection can be performed on the output from Google translate? The authoriship entropy contained within the original English text is likely to be carried through (at least partially) to the post-translated output.
Ask HN's are awesome but you don't usually get a PhD thesis in reply :-)
If so, perhaps it would be possible to use it to engineer countermeasures. In the example you used ("since" vs "because"), it would be fairly simple to alter the ratio.
Of course I don't know what more complex indicators you may be using, but I'm having a hard time imagining what would be easy to measure, but difficult to alter.
I guess you did not look at texts with multiple authors, or professionally edited texts? I am curious if a different editor, or publication house style, can be detected.