Time to Abolish "Statistical Significance"?
conversableeconomist.blogspot.com
conversableeconomist.blogspot.com
I agree that the focus on this metric negatively influences research outcomes, which extends to university structuring, but I'd like to hear your thoughts on how this extends to abuse of knowledge in general.
A nice way of saying "lying". Learning how to game the statistics (e.g. publish the 20th experiment that showed significance but fail to mention the other 19).
That is, if you have a theory about how a Gene relates to height on tomatoes, and you do a test, that test would show you you're likely on the wrong track if it falls below some p value, but the only thing it tells you by being above is that "there may be something here."
I think this is true for many fields with a replication crisis. The problem isn't statistical, the problem is no theory. If you have a functional theory there's all kinds of things you do to gain confidence in it, and mostly those will contribute to the ability to predict statistical results, but that is completely different on kind than sending out a survey and noting that question 2 and 6 are statistically correlated.
When a field thinks that the kind of early suggestive work like this is worth talking about, they should probably just talk about it in conferences and similar venues, rather than "publish" it where journalists will pick it up in a "science shows" story that 95% (lol) of the time turns out to be wrong.
In other words, I think it is fine that fields talk about early non-theory results -- that can be interesting for specialists to advance faster. "Publishing" this mostly-going-to-be-wrong stuff is leading to confusion among the public about what the scientific process demands and how trustworthy it is. That is not a good outcome in my opinion.
"We conclude, based on our review of the articles in this special issue and the broader literature, that it is time to stop using the term “statistically significant” entirely."
0.05 is way to stringent if you have 10 samples and way too lenient if you have 1 million samples.
But, by force of convention, everyone is using 0.05 (a value suggested by Fischer when basically all datasets were small) independently of their sample size and in a world were we are sometimes reaching dataset size that would have been inconceivable when the threshold was suggested.
Here is a good article to start to think about how one would select a p-value for an experiment : https://journals.plos.org/plosone/article?id=10.1371/journal...
But then you have to be able to calculate how costly is a type I and a type II error! That's seems a relatively straightforward question for a business (for example in A/B testing), but how do you measure that cost in academia?
I think this would only introduce confusion and another variable for p-hacking.
The strenght here is that you get rid of an arbitrary decision (p-value) and instead use quantities that can be measured and critiqued by a skeptical reviewer.
But, in my experience with small data, the impact of the size of the dataset dwarfed the impact of the cost/probability.
Not a magic number, just a threshold with consensus.
There is consensus that a p-value of 0.95 is reasonable. No one is arguing against that parameter in particular.
The argument is rather against the idea of statistical significance itself, because it's relatively easy to cheat if you are dishonest. Setting p to 0.01 won't change that.
Finally, and this is my personal opinion, there is "consensus" for abolishing statistical significance the same way there's "consensus" that Python is the best programming language. Scientists are more or less in agreement that it could be better, but "mildly displeased scientists call for unclear improvement" is not as good a headline.
See e.g. Redefine statistical significance. The authors list is practically a who's who https://psyarxiv.com/mky9j
>One Sentence Summary: We propose to change the default P-value threshold for statistical significance for claims of new discoveries from 0.05 to 0.005.
Furthermore, a consensus today does not need to hold up indefinitely.
Historical consensus, then.
Imagine you've done a study of something else and you accidentally discover a (properly FDR controlled) correlation between some drug and heart attacks, significant at 0.03 and with no big contrary prior.
Should you publish or wait 5 years for a follow-up study to complete?
What if three different groups stumble across this same finding? Should they all publish and maybe someone will realize 'shit this drug is killing people' or should they all wait to hit some other higher standard of significance?
The point is theres always a balance between specificity and sensitivity, a tradeoff in terms of costs.
I'd personally be happy with keeping 0.05 as the threshold for 'probably something here'. The real issue are about publication bias, incentives, naive interpretation of published work 'heres one study in Nature so it must be true'. I don't see any other purely statistical change, Bayesian (essentially has all the same problems, aside from taking into account priors) or whatever, that will solve these in any way that won't come without unacceptable sensitivity cost.
https://arxiv.org/abs/1904.06605
Victor Coscrato, Luís Gustavo Esteves, Rafael Izbicki, Rafael Bassi Stern — Interpretable hypothesis tests (2019)
Abstract:
Although hypothesis tests play a prominent role in Science, their interpretation can be challenging. Three issues are (i) the difficulty in making an assertive decision based on the output of an hypothesis test, (ii) the logical contradictions that occur in multiple hypothesis testing, and (iii) the possible lack of practical importance when rejecting a precise hypothesis. These issues can be addressed through the use of agnostic tests and pragmatic hypotheses.
Note that this enables acquiring one of the Holy Grail of Statistics, namely, controlling Type I & II errors simultaneously.
It is very different to confirm a prediction (i.e. to look for a particle with a precisely predicted mass), than to fish for some unexpected signal in your data.
In some cases Economics could do the same: Looking for an effect in any age range could be post-processed to take into account that you are looking into many age groups.
There are two alternatives to the current methodology:
* Remove significance requirement for publishing
* Adopt another statistical measure like Bayesian stats
There's a potential for abuse (as always), but perhaps you could curb it by mandating that you can only use the same co-submitter once. The other (likely bigger) problem would be figuring out how to get funding for it.
Full disclosure, I've never worked in a academia before, so please take this all with a grain of salt.