Finding Sociopaths on Facebook
schneier.com
schneier.com
> The problem isn't just that such a system is wrong, it's that the mathematics of testing makes this sort of thing pretty ineffective in practice. It's called the "base rate fallacy." Suppose you have a test that's 90% accurate in identifying both sociopaths and non-sociopaths. If you assume that 4% of people are sociopaths, then the chance of someone who tests positive actually being a sociopath is 26%. (For every thousand people tested, 90% of the 40 sociopaths will test positive, but so will 10% of the 960 non-sociopaths.) You have postulate a test with an amazing 99% accuracy -- only a 1% false positive rate -- even to have an 80% chance of someone testing positive actually being a sociopath.
Interestingly here he uses percentages to describe base rates and risk. Gerd Gigerenzer has a nice book, Reckoning with Risk, where he explains with many examples the problems of this approach. Gerd asks people to use real numbers instead, which are much easier to understand for most people.
Thus, Schneier's example becomes:
> Out of 1,000 people about 40 of will be sociopaths. You have a test that will tell you if someone is, or is not, a sociopath. The test will be correct 9 times out of 10. Bob has taken the test, and has been identified as a possible sociopath. The chance that Bob is actually a sociopath are actually about 1 in 4. This is because the test will tell you that 36 of the 40 sociopaths are sociopaths, but it will also incorrectly tell you that 96 non-sociopaths are sociopaths.
My writing is lousy, and other people will be able to clean this up, but even with my poor writing style it's easier for most people to follow and understand than the percentages.
This is alarmingly important when you're making a health decision - "Should I remove my breasts to reduce my risk of breast cancer?" for example.
(http://www.amazon.com/Reckoning-Risk-Learning-Live-Uncertain...)
EDIT: I use "sociopath" because it's in the source article. I agree with NNQ that it's very troubling to bandy around diagnostic labels like this, and deem people to be dangerous, just because of a tentative probabilistic diagnosis.
Honest question, what would happen to the other three quarters?
What Schneier is missing is that while you can't ID people that well from a single test, you can apply a bunch of them. In his example, one test improves the probability of correctly ID a sociopath from 4% to 24%. Apply another, different test of similar efficacy to that result set and you'll have a population of 21 true positives, and 8 or 9 false positives, increasing the probabiliy of a successful ID from 25% to ~70%. Sure, there's no single test that will give you reliable answers, but so what? It's OK to use a multi-pronged solution.
"Facebook records reveal convicted killer wrote 13-word post 5 years ago - red flag was raised - why was nothing done?"
Waiting until someone actually commits a crime will stop people being persecuted for a coincidental similarity of their behavior to that of a terrorist, sociopath or mime artist.
All because people can't understand the example in question, which appears in the first few chapters of most introduction to statistics books. And while all that money is being spent on useless checks the 9/11 terrorists, who the agencies were warned about, and the Boston bombers, who the agencies were ALSO warned about, are not followed up on because human and other resources are being spent on mass surveillance.
The internet is well known as a negative influence on certain people, but couldn't it be having a positive effect that is harder to measure and more an unintended side effect.
The real goal is building social/political systems that are robust and have checks and balances so that they cannot be perverted by special interests and are accessible to those who need them (child abuse support lines are a good example). Anything where a group intervenes on behalf of an individual is prone to disaster.
Are search engines and archives going to all willingly 'forget' that data when you 'just leave' Facebook? Are they going to not aggregate and correlate it to any new service you join?
This is one of the huge points of criticism of Real-Name-required services: a person can never escape an unjust judgment of such communities, due the long memory of the internet.
... would be pointless, as a perfect world would have no concept of "trouble".
It is not a statement about the world in general.
Even 0.4% of the population that get tested and are incorrectly "proved" innocent of being a terrorist amounts to more than a million undetectable terrorists in the US alone.
You're saying that you would be happy to join 74 other non-terrorists (i.e. law abiding citizens) plus 25 actual terrorists and be taken off to Guantanamo Bay indefinitely ?
You're really sure about that being a Good Thing for law enforcement ?
But today try to ask a few people around you, and see what they say.
> Suppose you have a test that's 90% accurate in identifying people who have a disease, and 90% accurate in identifying people who do not have the disease. Assume that 4% of people have this disease. Hypothetical_Bob is tested, and the test says that he has the disease. What are the chances that Bob actually does have the disease?
Lots of people - smart people too! - struggle with this. Even if you give them pencil and paper and let them doodle around they will often give you an incorrect number. And most of them will be surprised if you tell them it's as low as 26%.
I think my point is slightly orthogonal since I misunderstood you; if you tell someone that something is "10%" they will think "that is pretty bad" whereas "1 in 10" is more likely to get a "hey, that's not too shabby" response. Percentages sound "worse" than numbers, even when they are the same (at least to me). Perhaps because they are harder to reason with?
Suppose you have a test that's 90% accurate in identifying both people with X and people without X. If you assume that 4% of people are people with X and you're told that you test positive for someone who has X.
Do you really find it easy to arrive at your actual chance (26%) of having X? Let's not forget that most people on HN are at the smarter end of the bell curve. It'd be interesting to see the results of a large scale study about answers to questions like this.
When I said I found his version clearer, I meant between the two versions originally given. The one you've just added is of course less clear because unlike the other two, it doesn't point out the issue.
BTW: But maybe it is clearer, for calculation rather than understanding, because I get 27.(27)%, not 26%... https://www.google.com/search?q=%28.9*.04%29/%28.9*.04%2B.1*...
true positives .90 * .04 = 0.036 false positives .10 * .96 = 0.096
total positives 0.132
positives that are true positives 0.036 / 0.132 = 0.2727...
i had to think about the calculation as i was doing it wasnt automatic even though it was just multiplication, but I think the difficulty is more to do with the fact that you have to use some relative of bayesian probability not really the fact that you had to deal with percentages
Even wildly inaccurate tests can be useful. Imagine that driving drunk will result in a accident 10% of the time. 90% of the time, however, a driver will make it home safely. This is a wildly inaccurate predictor, but it still critical to know someone's blood alcohol level before giving them their keys.
One must simply be aware of the uncertainty involved in any test, and treat test results as probabilistic signal, not as proof.
Look at the numbers of people getting very serious medical treatment because they, and their clinicians, have not understood the numbers.
If doctor (well educated intelligent person) cannot get this right I'm scared that labelling someone as "POTENTIAL TERRORIST" on the basis of a 1 in X possibility is going to have disastrous consequences.
Of course, discovery does not require treatment, so you could test, find a positive, and not do anything. But EBM people tend to view idealized responses with suspicion, and some argue that taken in real-world conditions, not administering certain tests, or administering them in more restricted situations, or at least not recommending them as the default, would improve aggregate outcomes (and the data seems to support that). They would then restrict the tests to cases where testing statistically improves outcomes.
Same logic applies in other situations.
The field is laughable, really. It's sad that people making important decisions in the medical and pharmaceutical fields cannot correctly interpret a confusion matrix, or if they can they are corrupt and decide to ignore it anyway.
You really don't want to apply the same standards to law enforcement, as they are appalling.
EDIT: to the downvoters, I'd suggest doing some reading. In particular, check on the existence of any verifiable testing that proves any semblance of correlation between HIV and AIDS. The standard is appalling for any scientist who's willing to check on the numbers. In general, the standard in medical science is rather low compared to say physics, but in this case it's astonishingly low considering the accolades Montagnier got for his research. I understand it's easy to dismiss this as conspiranoia but it really is not. There are dubious claims passed as truths with very specific interests behind. Not making an assertion either way on the HIV<->AIDS relationship, just pointing out that the standard of the research is shocking and the statistics skills of the people involved are incredibly poor, if not intentionally corrupt.
Using condoms ever killed anybody. Stay save. And if you personally refuse therapy, ok. But HIV does cause AIDS and has a one hundred percent lethality. In case you affected use the remaining time.
The fact that AIDS is virus-induced, that remains unproven. However, it's pandered as such and a number of companies are cashing big on antivirals and antiretrovirals, in many cases worsening the patient's health. To claim something is proven without proof, that is pathetic.
And this is not the only scam induced by big pharma. There are many other perverse effects stemming directly from the fact that pharmaceuticals cash big on chronic illnesses, giving them an incentive to research not in cures, but in long lasting treatments.
Feel free to downvote away though. It's important and downvotes might actually get more people to read it.
Even if you get a system good enough to overcome the base rate problem, you'll only end up labeling a bunch of mostly harmless people. Think a about a hypothetical uber-villain that would want to recruit children or teenagers, brainwash them and turn them into assassins or other kind of agents. He may find out that people with some sociopathic traits are better candidates for this, so he will target them. Now think the uber-villain is, uhm... (working for) your government :)
...not to mention the mislabeling of people with atypical social interaction patterns, like ones with mild/pseudo aspies which combined with the base rate fallacy brings serious mislabeling.
I'm sure this kind of electronic-psycho-profiling is already in use, and I even think it may have interesting side-benefits, like cool work being done in AI research for use in this (no better way to start "humanizing" and AI than to have it model human personalities and predict their actions), but there's tons of things that can go south with it for lots of innocent people that just happen to be "different" (like most people who end up making breakthrough discoveries or world changing inventions, you know...).
It can be argued that the lack of empathy can lead to a more violent behavior, but violent acts comming from deep personal relationship are plenty a dozen as well, who knows.
The DSM doesn't even list psychopathy as a diagnosis any more, only antisocial personality disorder which requires a history of, and we might as well just quote the DSM "... a pervasive pattern of disregard for, and violation of, the rights of others..." So yes, a diagnosis of ASPD does generally indicate someone who could be considered dangerous to others.
Now if we're talking about someone with psychopathic traits that scores high on a Hare, then no it does not necessarily indicate dangerous behavior. However, a Hare is still important when dealing with criminals as you do not want to give a psychopath treatment as it simply makes them better criminals.
What do you mean by "treatment" in this context? I thought that caught psychopathic criminals can end up in criminally insane facilities and such in most developed countries and somehow treated or at least attempt to do so.
I suggest you research "natural killers" from 1946 onwards (famous paper starts the ball rolling there - Combat Neuroses / Fatigue is a tip). Their percentage increase in producing effective killers is impressive. I won't link to specific papers, since it's outside the remit of H.N., however a lot of hard science & tech has been brought to bear on the issue.
The flip-side of this, the mirror to a sociopath (let's call it "the empathetically linked / driven") makes just as an effective killer for the record. If not better.
Tl;dr
You're about 70 years too late to pioneer this field. However, in looking @ the current trend to map Autism onto AI networks as a model, I'd suggest you'd probably have more joy looking at the empathetically linked individuals to see how their strong network connections / protective instincts are harnessed if I were to build a neural Map (weak) AI.
My immediate thought is that if the base rate fallacy were part of their education, society would be better off, but legislators still have to play to their constituents and it's hard to have hope in that sphere.
So you get a quote through their app, which uses your likes, your social network, and maybe some NLP on your posts to decide if you qualify for a lower quote.
Insurance people in the UK tell me this would be very useful, but probably isn't possible from a legal perspective, but there will be other countries where it is.
An effect of this would be to penalize people who don't use social networks.
"Even bad men love their mommas." - Ben Wade (Russell Crowe), 3:10 to Yuma.
The problem is that there is always enough people who are not normal (there is a great book about it by Erich Fromm: The Anatomy of Human Destructiveness) and it's possible to 'make' people into something not normal in psychological sense - i.e. suppress their empathy for certain group of people by some war trauma or conditioning/brainwashing.
BTW, I think that it's a myth that even bad men love their mommas. It depends how you define love. From what I remember based on one psychopath that I know really well he claims that he loves his mother and maybe he even thinks he feels something like that but he acts in such a way that his mother often gets hurt by his actions and he acts with complete disregard of that. Not love in my book.
Who is going to label the data with ground truth? Clinicians who "know it when they see it"? What is the ground truth that the classifier is going to train on?
If you're going to do an unsupervised classifier (eg clustering) who is going to label the clusters? What is going to keep the data from turning into uncorrelated mush?
So I question the basis on which Mr. Adams -- whose works some might consider a sociopathic attack on American business practices, and by extension, capitalism -- thinks content analysis is a good idea.
http://blogs.scientificamerican.com/cross-check/2013/05/04/p...
If that'd be the case, would facebook just be some sort of crime catching tool ?
I wonder though, is sociopathy a crime, or some trait that will turn people into criminal ?
> I wouldn’t be surprised if some of the Facebook members who are most active and have the most ‘friends’ are sociopaths.
I doubt the people you are talking about can be criminal.
Crime and psychology, sigh
> I don't think criminal hang around facebook.
Are you sure about that?
However, the common burglar is probably not a sociopath, but rather a poor drug addict or something.
When you look them up in the DSM (the big book of psychological disorders), it will seem that way. However, you can be antisocial or possess sociopathic traits without having a disorder. In general dictionaries, ’antisocial’ is synonymous to ’unsociable’. That’s how it’s most often used and that’s how I meant it.
> Perhaps you're thinking of people who are non-social?
I had never heard of that word. After Duckducking it, it seems to me that no one uses it. If I were to use another word, I’d probably go for ’asocial’.
They are lifes hackers.