“Should this even be released?” Deep learning tool that may be used for doxxing
github.com
github.com
[1] https://www.aclweb.org/anthology/E/E99/E99-1021.pdf [2] http://ceur-ws.org/Vol-1391/126-CR.pdf [3] https://www.aaai.org/ocs/index.php/FLAIRS/FLAIRS13/paper/vie... [4] http://www.cs.waikato.ac.nz/ml/weka/ [5] http://ntv.ifmo.ru/en/article/15185/kompyuternaya_kriminalis...
I mean sure, there are various 'immoral' uses for it (like doxxing), but there are also many good ones. Such as:
1. Working out who wrote a bunch of anonymous reviews on Amazon or other such sites, which could be used to stop fake reviews. You actually mention this usage in your article.
2. Being able to identify troublemakers in a community (such as a forum or a social networking site). I'm sure a lot of administrators would love to know if that suspicious looking new guy is the alias of a banned troll from a few weeks back (posting through a proxy server).
3. Literary analysis, like working out who wrote many anonymous works of fiction. Or as someone said below, determining which parts of Shakespeare plays were actually written by Shakespeare.
4. Crime solving. If it works anywhere near as well as you say, it could theoretically help unmask the Zodiac Killer, or perhaps even Jack the Ripper (if any of those letters were real).
All the uses above would be a net positive for humanity, and would be great possible uses for a deep learning tool like this.
Don't let the worries about its usage by 'bad' people overshadow the good you can do by releasing it.
Orwellian.
As in, flames everyone to a crisp, posts as much porn as possible, tries to incite a civil war between a few members that might not be on good terms with each other or the staff and registers hundreds of accounts, some of which stay semi dormant until they strike?
Because that can happen very easily online, especially if you get the ire of someone with a lot of free time and very few morals. Or if your site ends up at war with a troll site/gets raided by 4chan.
Do you avoid the hassle now, or wait until the situation blows up and half the site is now in the middle of it?
The idea that this tool would be useful for community management is terrible.
I'm a moderator of multiple online spaces.
A few months ago, in one of them, a user got too heated and started flinging insults at someone else. As was standard policy for the place where it was occurring, I issued the user a ban of a few days (enforced cooling-off) and pointed to our guidelines on how to behave.
This user then proceeded, over a period of months, to continually harass me, send me increasingly graphic threats, and try to track me down in real life.
Pray tell, how exactly would you go about "converting" such a person to be productive? I come to you since you are apparently quite the expert on it, or else you wouldn't be giving out advice to just "convert" people.
I voted to release the tool, but you've set up this false scenario with the intention of knocking it down easily and discrediting the opposing argument.
I don't know how well that would work for high-traffic forums of today, but it can scale pretty easily by employing many moderators.
Ask anyone who's ever lost an ebay or a google account to an algorithmic burp and was essentially banned for life without appeal or even human oversight.
1 and 2 could both be a bet negative for humanity just as easily as they could be a positive.
Or maybe an odd case where it turns out the author of a book or creator of a product finds out someone they know in real life left the negative review and physically attacks them or something.
Unfortunately if the tech exists, it will be released, but I don't think the positives outweigh the negatives here. The chilling effects in terms of comments alone would be bad.
In fact I believe this tool is even more worrisome, because there are a very large number of non-tech savvy people who express their dissonant opinions simply under the mask imparted by internet anonymity. I imagine most everyone here on HN have at least once made an anonymous account to post a comment somewhere that they would rather not have tied to their identity. This is an avenue that is necessary for the preservation of free speech. Remember that things like treating black people as equals, giving women the right to vote, gay rights, etc were and in some ways still are taboo subjects that bring the wrath and ire of the power du jure.
All that said, I still don't see any reason why such a tool shouldn't be released. Why? Even though I believe the tool to be harmful, it's better to know that it exists, know its capabilities, and most importantly know how it works. It'll end up in the hands of the wrong people anyway, so it's better to at least get it into the hands of the right people who can possible do something to combat it.
This is a really important point. Even if you don't like the possible uses of the tool, it's either "release it now and make it possible to defend against" or "don't release it, and hope the likes of the NSA don't develop their own version".
>>>>>>>>>>>>>>>>>>>>>>> LOGGING STACK TRACE <<<<<<<<<<<<<<<<<<<<<<<<<
java.lang.IndexOutOfBoundsException: Index: 0, Size: 0
java.util.ArrayList.rangeCheck(Unknown Source)
java.util.ArrayList.get(Unknown Source)
edu.drexel.psal.anonymouth.utils.FunctionWords.run(FunctionWords.java:60)
edu.drexel.psal.anonymouth.engine.DocumentProcessor.processDocuments(DocumentProcessor.java:140)
edu.drexel.psal.anonymouth.engine.DocumentProcessor.access$000(DocumentProcessor.java:39)
edu.drexel.psal.anonymouth.engine.DocumentProcessor$1.doInBackground(DocumentProcessor.java:70)
edu.drexel.psal.anonymouth.engine.DocumentProcessor$1.doInBackground(DocumentProcessor.java:67)
javax.swing.SwingWorker$1.call(Unknown Source)
java.util.concurrent.FutureTask.run(Unknown Source)
javax.swing.SwingWorker.run(Unknown Source)
java.util.concurrent.ThreadPoolExecutor.runWorker(Unknown Source)
java.util.concurrent.ThreadPoolExecutor$Worker.run(Unknown Source)
java.lang.Thread.run(Unknown Source)
( ErrorHandler ) - Fatal error encountered, will exit now...
It's too bad the flow isn't more along the lines of "Give me some docs from one author, and the other document you want to test. Okay! Here is your result!" anonymouth/src/edu/drexel/psal/resources/koppel_function_words.txt
And down the rabbithole we go; after copying the file to src/jsan_resources we now get a new error: >>>>>>>>>>>>>>>>>>>>>>> LOGGING STACK TRACE <<<<<<<<<<<<<<<<<<<<<<<<<
java.lang.NullPointerException
edu.drexel.psal.anonymouth.engine.InstanceConstructor.getAttributes(InstanceConstructor.java:182)
edu.drexel.psal.anonymouth.engine.InstanceConstructor.runInstanceBuilder(InstanceConstructor.java:121)
...
>>>>>>>>>>>>>>>>>>>>>>> LOGGING STACK TRACE <<<<<<<<<<<<<<<<<<<<<<<<<
java.lang.NullPointerException
edu.drexel.psal.anonymouth.utils.TaggedDocument.makeAndTagSentences(TaggedDocument.java:380)
edu.drexel.psal.anonymouth.engine.DocumentProcessor.processDocuments(DocumentProcessor.java:184)
...
:)I can only assume it was tested on a synthetic dataset perhaps.
Also I'm wondering how many unique users are present in the dataset, along with the volume of content for each user.
A great test set for anyone trying to do this: look at the Scott Adams sock puppet controversy on Metafilter [1] and see if you can train something on his public writing to match the "PlannedChaos" commenter's posts and Adams' own tweets. It is probably the closest you could get to a "pure" training set in the sense that presumably Adams didn't think he'd get caught. (And if he did, and therefore did alter his stylometrics, then it's even a better challenge.)
[1] http://www.adweek.com/galleycat/scott-adams-caught-defending...
For example, suppose someone is revealing on Reddit details about some business dealing of yours that should have only been known by people who are under NDAs. If you intersect the set of 50000 Redditors returned by the deanonymizing tool with the set of people under your NDA, and that intersection is not empty, then the leaker is probably one of the ones in the intersection.
Can you imagine how much it would suck if you woke up the next morning and the entire internet is convinced that you're Satoshi Nakamoto or a pedophile due to a false positive from this program? There is no due process and no chance of appeal; your social life is simply over at that point. All because of a 5% chance of a false positive.
The limit is in attention paid -- the public has a limited capacity to absorb information, and there are a few hundred, perhaps a thousand or so "top celebrities" at any one time.
And some of those can attain a highly significant level of immunity to criticism. Ronald Reagan's presidency was the most scandal-prone in recent memory, and yet his moniker was "the Teflon president". William Jefferson Clinton took far more flack for far less, and Barak Obama takes the hit for complete fabrications. Meantime, a major party presidential candidate advocates overt violence to protesters and various other views ... and is only embraced all the more strongly by his supporters.
The dynamics of this are odd.
That said, I'm not sure the genie can be rebottled. It's an area of privacy in which law rather than technology must be applied. Including, say, exceptionally fierce penalties for misuse.
In cases like this, I feel erring on the side of being less explicit tends to help. Leave just enough out to let everyone read between the lines, and put the pieces together.
Don't say it's a "Machine learning algorithm to connect anonymous accounts to real names", say it's a 'speech pattern analyzer', or say it 'allows comparison of speech patterns for likelihood of same author'.
The people who need this for evil purposes will develop it whether it's released in open source or not.
This sounds like a beginner who created a dataset, with a flawed metric. And is now going around claiming 95% accuracy, using "Deep" learning. And equally clueless commenters are hyping it up.
Why stop at claiming 95%, hell even I can create a "dataset" and a "deep learning" algorithm and get 99.9%.
I am not discounting that there are legitimate stylometric analysis methods, which have been peer reviewed. But please lets not hype "Deep learning for doxxing". This just sullies the real progress being made in deep learning.
That said, identification may be a more tractable problem if you have a limited population, additional metadata for features, and normalized writing samples (comparing anonymous reviews to identified reviews, or within community posts, as opposed to trying to compare a set of anonymous tweets to an identifiable dissertation).
Generally in supervised machine learning a claim of X% accuracy means that when tested on a large dataset for which the correct result is known and that was not part of the training dataset or validation dataset, it classified X% of that dataset correctly.
Typically you gather a big dataset of labeled data and then split it randomly into training, validation, and test sets. A 50/25/25 split is common. If the learning approach you are using does not need a validation set, then 70/30 training/test is common.
How reliable such an accuracy estimate is depends on how well your dataset matches the characteristics of the datasets people will be using your trained system on. His 95% accuracy report is probably reasonably reliable when his software is used on anonymous posts on the forums where he gathered his datasets. It would probably be less reliable looking at anonymous posts on, say, a bagpipe maker's forum.
However even in supervised learning, accuracy is only used in very limited cases such as multi class classification. For a whole bunch of problems including the one being discussed its a poor and in some cases a biased metric. E.g. consider a heavily unbalanced problem 99% positive labels. By predicting all instances with majority label its possible to get 99% accuracy. There are several better metrics, False Accept rates, Precision Recall curves etc.
Without knowing how the dataset was collected, did the "username" leaked into the dataset, etc. its impossible to evaluate such outlandish claims.
The whole moral and ethical debate is non-sequitur, and harms legitimate deep learning research.
...but what the most effective classifiers do is memorize the names and signature blocks of people who posted in each newsgroup.