Machine learning of neural representations of emotion identifies suicidal youth
methodsman.com
methodsman.com
Anyone with a Nature subscription want to check whether they simply trained their discriminator and then used it on the same data set? There's no mention in the abstract of testing it against a fresh control set and that's not promising.
https://www.nature.com/articles/s41562-017-0234-y?error=cook...
"A Gaussian Naive Bayes (GNB) classifier trained on the data of 33 out of 34 participants predicted the group membership of the remaining participant with a high accuracy of 0.91 (P<0.000001), correctly identifying 15 of the 17 suicidal participants and 16 of the 17 controls"
Sounds like they overfit their cross validation score and reported that. The data is actually available here though:
We have no idea how this result would work on a new data point that has not been used in training.
It's bad statistics, bad data science.
> One round of cross-validation involves partitioning a sample of data into complementary subsets, performing the analysis on one subset (called the training set), and validating the analysis on the other subset (called the validation set or testing set). To reduce variability, multiple rounds of cross-validation are performed using different partitions, and the validation results are combined (e.g. averaged) over the rounds to estimate a final predictive model.
https://en.wikipedia.org/wiki/Cross-validation_(statistics)
You might object that it's difficult to achieve the level of inter-fold isolation required to make the technique sound, and indeed if you search the comments there's some question as to whether or not there was an information leak in their feature selection process. In that sense calling it "bad statistics, bad data science" might be reasonable, but it's also a powerful technique, so I don't think it's reasonable to dismiss out of hand without being more specific.
Looking at it either from a machine learning or statistical point of view, using such a small sample is problematic.
This is the chronic issue with fMRI studies, since administering an fMRI is extremely expensive, and has led to some very difficult to reproduce results in the field.
Example: you believe a newly found plant species is toxic. You give it to 17 "grad students volunteers", while giving a placebo to 17 others. All in the first group die aa gruesome death within 20 hours. None of the others do.
Result: yes significance. (also: tenure!)
I'm not saying that this study is significant (the statistics seem to be slightly beyond my event horizon), and your criticism also stops short of an outright dismissal of the research. But sample size alone makes for a bad measure of quality. Yes, even p-values are better.
Effect size is very important in this. To continue your grad student murder example, it's completely trivial to determine which plant a student was given, based on whether they are dead or not. It becomes trickier if you measured something a bit less cut-and-dry, such as the incidence of headaches, or variance in a few voxels of a noisy MRI.
(Consider what happens to people so-diagnosed as suicidal when in fact they are not (false positives). Involuntary psychiatric imprisonment is a terrible thing if it isn't absolutely necessary.)
[0]: https://www.healthline.com/health/depression/facts-statistic...
IANAStatistician, but let’s consider the system is right 91% of the time and we try to detect those 7% you mentioned. Let’s take 1000 people. 70 people are depressive and 930 aren’t. Out of those, 700.91=63 will be correctly classified as depressive by the system and 9300.91=846 will be correctly classified as non-depressive.
That leaves us with 63 positives, 846 negatives, 7 false negatives and 84 false positives. False positives largely outnumber false negatives, but they also outnumber the true positives.
(if a statistician read this, please correct me if I’m wrong)
Will it be that accuracy actually means AUC?
Will it be that they are reporting predictive skill on the training data?
That's because the measured difference between the groups would be lower (because the real difference would be lower if the groups are more alike than you think).
Say you're testing a drug that's supposed to make people taller. You don't know it yet, but it really does make everyone grow 10cm overnight. You give it to half of your volunteers, and the other half gets placebo. The next day you find that the first group grew by 10cm compared to the control.
Now say your grad student messed up and half of the control group also got the real thing instead of placebo. Those also grew by 10cm, making the average in the control group 5cm, and your treatment group's effect is suddenly lower.
[...]
The features used by the classifier to characterize a participant consisted of a vector of activation levels for several (discriminating) concepts in a set of (discriminating) brain locations. To determine how many and which concepts were most discriminating between ideators and controls, a reiterative procedure analogous to stepwise regression was used, first finding the single most discriminating concept and then the second most discriminating concept, reiterating until the next step reduced the accuracy. A similar procedure was used to determine the most discriminating locations (clusters)." https://www.nature.com/articles/s41562-017-0234-y
The winner is #3: data leakage leading them to use predictive skill on the training data.
How so? The data used as validation in one fold would be used to determine features in the next...
"To identify the most discriminating concepts, a reiterative procedure analogous to stepwise regression was performed. In the first iteration, the group classification was performed using only one concept at a time, determining which single concept of the 30 resulted in the highest classification accuracy. In the second iteration, the classification was performed using pairs of concepts, namely the single concept that produced the highest accuracy in the first iteration as well as each of the 29 other concepts. All pairs that produced at least as high an accuracy as achieved on the previous iteration, were explored in the third iteration, where triplets of concepts were used, namely the pairs that produced the highest accuracy in the previous iteration, plus each of the remaining 28 concepts. Such stepwise addition of discriminating concepts continued until adding any one of the remaining concepts resulted in a decrease in accuracy. An analogous procedure identified the most discriminating locations."
But I still think even in your case they are doing:
train: abc; val: d -> score1/ features0 -> features1
train: abd; val: c -> score2/ features1 -> features2
...etc
score2/features1 would all contain info from c, etc.IANAStatistician, but this seems like a trash result.
"The features used by the classifier to characterize a participant consisted of a vector of activation levels for several (discriminating) concepts in a set of (discriminating) brain locations. To determine how many and which concepts were most discriminating between ideators and controls, a reiterative procedure analogous to stepwise regression was used, first finding the single most discriminating concept and then the second most discriminating concept, reiterating until the next step reduced the accuracy. A similar procedure was used to determine the most discriminating locations (clusters)."
The features were chosen using the same data as used to assess predictive skill.
Can you provide pseudocode consistent with what they described (in the post you responding to) that wouldn't lead to leakage? I can't see it.
To get the estimation variance down, you can repeat this for all possible choices of validation sample. That means, you start the feature selection process on the new training set over from scratch and obtain another risk estimate. If they kept the features selected earlier, that estimate would be "contaminated" and not independent, but if they correctly start over, the procedure is valid.
When we want to use these models, we run new/test data through all N=34 models in parallel and calculate a prediction from each. Then somehow these predictions need to be combined (one again an average, etc). This is the average of the predictions, not accuracies/whatever.
Where was the step combining these predictions present during the training? It seems your scheme necessarily calculates an accuracy based on a different process than needs to be applied to new data.
Of course you could build an ensemble model, but if you want to know the expected accuracy of doing that, you need to include the ensemble-building into your validation procedure. (Or use some theorem that lets you estimate the ensemble performance from that of individual models, if that is possible.)
Using which set of features? You have 34 different models with different features...
After deciding on features/hyperparameters (based on the overfit cv), you train the model on all the data used for cv at once. Then test the resulting model on a holdout set (that was not used for the cv). The accuracy on that holdout would then be the accuracy to report.
This sounds much like what you are describing, except you only do one cv and do not use it to decide anything. The cv is only to give an estimate of accuracy.
Is that correct? It does seem to legitimately avoid leakage. However, it seems impossible that an anything close to optimal feature generation process or the hyperparameters were known beforehand. Do you just use defaults here?
https://www.naturalblaze.com/2017/03/scandal-mri-brain-imagi...
[0] https://www.sciencealert.com/a-bug-in-fmri-software-could-in... [1] http://www.pnas.org/content/113/28/7900.abstract
Confusing the two would lead to the more unusual conclusion that suicidal ideation is associated with abnormal brain connectivity, while the authors are instead focusing on neuronal activity.
[0] i.e. you know with precision where in the brain activity occurred, but less precisely when it occurred in time
Diagnosis tools could mean faster access to treatment. Currently in the UK the waiting list for access to mental health treatment is on the range of two to three years. Transforming "suicidal ideation" from a "vague human-given diagnosis" to "tool-given diagnosis" makes it politically easier to push for that.
In any case, that's not going to happen based off a single study with 91% accuracy.
Tip: If you own a gun and are feeling suicidal, give it to a trusted person for safekeeping.
"A study by the Harvard School of Public Health of all 50 U.S. states reveals a powerful link between rates of firearm ownership and suicides. Based on a survey of American households conducted in 2002, HSPH Assistant Professor of Health Policy and Management Matthew Miller, Research Associate Deborah Azrael, and colleagues at the School’s Injury Control Research Center (ICRC), found that in states where guns were prevalent—as in Wyoming, where 63 percent of households reported owning guns—rates of suicide were higher. The inverse was also true: where gun ownership was less common, suicide rates were also lower."
https://www.hsph.harvard.edu/news/magazine/guns-and-suicide/
Everybody points to the Australian example, where suicides declined after the 1996 gun control legislation. But unemployment in Australia peaked in 1995 and declined precipitously afterward until 2009.
Given everything we know about suicide rates in other countries, and about changes in suicide rates domestically (e.g. recent increase as gun ownership goes _down_[1]), it would very odd if gun ownership was a root cause of suicide.
That said, in a country with a strong gun culture like the U.S., I would totally expect a generational dip in suicides if we substantially removed access to guns. But then I'd expect it to normalize when suicidal individuals became more comfortable with other methods. Just like with mass shootings, there's a strong imitation effect. Take away the model that people imitate and it might be awhile until there's a regression to the mean.
Even so, that's still reasonable justification for limiting access to guns--saving tens of thousands of individuals. I'm not sure I'd agree with such a policy prescription because of the insane gun politics, but it's quite defensible from a public health perspective.
[1] Number of guns have increased but they're concentrated in fewer households.
Tl;dr?
* Many suicide attempts occur with little planning during a short-term crisis.
* Intent isn’t all that determines whether an attempter lives or dies; means also matter.
* 90% of attempters who survive do NOT go on to die by suicide later.
* Access to firearms is a risk factor for suicide.
* Firearms used in youth suicide usually belong to a parent.
* Reducing access to lethal means saves lives.
Suicide is an epiphenomenon of larger socio-economic issues. Among OECD countries the U.S. comes in the middle of the pack. If guns were a causative factor, then considering how prevalent they are (by a ridiculous factor!) our rates should be much higher:
https://en.wikipedia.org/wiki/Suicide_in_the_United_States#/... http://foreignpolicy.com/2013/05/03/how-does-americas-suicid... http://www.oecd-ilibrary.org/sites/health_glance-2011-en/01/...
Any correlation between guns and suicide in the U.S. is easily understood in terms of modeling--guns are how Americans kill themselves. Take away the guns and, yes, there'll be a dip in suicide rates, until Americans learn how people kill themselves elsewhere around the world. Heck, they're already learning that with opioids.
Let's go back to what I said about Australia. I claimed that the change in Australian suicide rates is better understood in terms of the unemployment rate. Now let's test that hypothesis. [... google google google ... ] Here we go:
Suicide has reached a 10-year high in Australia as 3027
people killed themselves last year, the largest cause of
death among 15 to 44-year-olds.
Last year, 12.6 people in every 100,000 killed themselves
compared to 12 the year before, 11.4 in 2012 and a low of
10.4 in 2006.
-- http://www.theaustralian.com.au/news/nation/suicide-rate-in-australia-reaches-10year-high/news-story/cb5d8384aadb571778775bda236f3c35
What was the rate in 1995, the year before the gun control law? 13.0. (https://www.aph.gov.au/About_Parliament/Parliamentary_Depart...)So suicides were at their lowest the very same year that unemployment was at its lowest (2006)? Check! And they rose as unemployment rose? Check! To the point where they're back at the pre-law level? Check!
I won't deny that there's some nuance here that we can tease out, but if you read the actual papers that link guns to suicide, they do a much worse job at nuance. In fact, not a single one of the papers I've read even considered the unemployment rate. Which is patently bad science.
Correlation is not causation. All the gun suicide papers do is point out specious correlations. But you don't need a degree in statistics to know this. And you don't need a science degree to be able to see the gargantuan holes in these arguments--that the correlations have simpler explanations.
I'm not denying that gun control could appreciably effect suicide rates. Imitation and modeling have huge effects--much more well-established than the supposed gun effect. So huge that even news media abstain from reporting suicides, especially methods of suicide. Saving thousands of lives with gun control, even if it's an ephemeral gain, is an absolute benefit that's worth debating about. But let's just be honest about this stuff.
It does seem like you agree in your last paragraph though though... gun control would lower suicide rates. People would try other methods which have higher failure rates, people's lives would be saved. Not all of them, but an appreciable amount.
Lots of Europe has a higher suicide rate than the US. Within the US, as the other response said, gun ownership is confounded by all sorts of other factors.
(It's also not clear that even if they have suicidal ideation that they're not entitled to gun rights, but I don't feel like having that debate today.)
I don't know about other states but it's likely any that require background checks for private transfers ("closing the gunshow loophole") would make this behavior illegal.
> Avoid unrelated controversies and generic tangents.