Is it time to up the statistical standard for scientific results?
arstechnica.com
arstechnica.com
This screams systematic error and error propagation to me. It's possible that we don't need to up the p-value, we just need to make sure researchers aren't stupid and can properly account for all sources of errors. The problem is that's often an acquired skill over time, not something younger researchers typically think about, especially those who aren't multidisciplinary.
For example, you could have an extremely unlikely hypothesis (eg. "dice are controlled by alien telepathy"), test it, and still come out with p < 0.05 because of something much more plausible ("the dice were badly made and are biased"). This is known as the base rate fallacy; an extraordinary claim requires extraordinary proof. Moreover, to prove an extraordinary claim, you must show it is more likely not just compared to chance, but compared to every less extraordinary alternative, including "the experimenters are faking the data".
It's also entirely possible to get a result at p < 0.05 that makes the hypothesis you're testing less likely. Suppose you want to know how far away the nearest star is. The value in the textbook is 13.4 light-years, and you think it's really 12,000. You take some measurements, and get values of 19.6, 17.4, 20.1, 20.4 and 18.5.
Now, this is a significant result at p < 0.05 - if the star is 13.4 light-years away, then getting these numbers has a probability of less than 5%. However, these results completely rule out the hypothesis you're testing. The numbers you get are pretty unlikely if the real number is 13.4, but extraordinarily unlikely if the real number is 12,000, so this experiment makes the 13.4 number more credible. This kind of thing is why lots of researchers still believe in psychic powers - they keep testing for psychic powers, and keep getting results at p < 0.05, but don't notice their results make psychic powers less plausible. (Not joking - see http://commonsenseatheism.com/wp-content/uploads/2010/11/Wag... for a detailed explanation.)
This is an elementary mistake that every freshman statistics course warns against, and yet it's absolutely pervasive.
It's got priiiiooooooors.
The hypothesis in question is the null hypothesis which must be assumed true not false for the p-value to mean a 100*p% chance of seeing the observed data.
I can see why elementary mistakes are so pervasive.
The thing I fear is underexperienced multidisciplinary researchers coming up with convoluted procedures that put their hands and eyes on their work in so many places that small sources of error have compounded themselves and entrenched themselves into nooks and crannies.
Very few degree programs give scientists enough background in statistics and experimental rigor to prevent these problems.
Perhaps that's a little bit of projection, but I don't use any but the most rudimentary statistics in my publications and I don't make claims about significance.
I think your comment about them being less likely to be able to handle depth from a personality/self-habits point of view is really projecting and probably based on experiences with a few bad apples. My experience doesn't reflect that at all.
I feel like a lot of the problems with statistics come from overly insular journals and conferences where everyone does it wrong and nobody realizes that because they're all doing the same thing. I've seen this in some computer science fields.
People seriously just take standard deviations and then plug it into the formula to spit out p values. Without thinking about things like: "should the values be normally distributed? or log-normal?" "How does error propagate through this formula I'm using?" "Is the major source of error in the replicates that I'm using (versus something I might be normalizing to, like a mass measurement)?" "Am I in the linear range of my calibration standards?" "Is using these data quantitatively (versus qualitatively) an honest thing to do?"
Reproducibility by researchers who know no more about the experimental setup than what was published is a better way of guarding against this than upping the significance hurdle. Experienced researchers may have a long history of ignoring confounding factors - I am not sure maturity matters much. I think the best time to have your complacency about constructing statistical investogations torn apart is while you are doing your PhD.
Eating an apple might help a broken leg heal faster or slower. There is certainly an extremely small correlation, and with a sufficiently large number of controlled experiments, a result rejecting the null hypothesis with p < 0.05 will be found. The required number of experiments might be astronomically large, but if any correlation exists, it can be found with enough samples.
But without looking at the effect size, the result is useless. Even if eating an apple helps your broken leg heal 2 seconds faster on average, it is pointless to suggest this as medical advice.
If you're just feeling your way around in the dark, 2-sigma is a useful way to work, so we use that to guide exploration.
Why 2-sigma? Well, it's twice as big as 1-sigma.
Experiment didn't go well, but you need a more-impressive result? Use a 90% confidence interval instead of 95%.
Unfortunately, many studies -- particularly those in medicine -- are already conducted with samples that are too small to detect any effect you'd reasonably expect to see, because many researchers do not calculate in advance what sample size would be required. This has interesting paradoxical effects: the only published studies are those that overestimate the size of the true effect.
http://www.refsmmat.com/statistics/power.html http://www.refsmmat.com/statistics/regression.html#truth-inf...
So there's a tradeoff. Do you want to eliminate false positives at the cost of more false negatives? It's a difficult balance. I suspect there are many areas where poor statistical practice can be remedied to produce better results without greater expense.
As I scientist who has worked with both rich, and poor datasets from standard and invented datatypes, I'd just say, "let me see the data".
No it isn't [1].
The best definition I know is
> The P value is defined as the probability, under the assumption of no effect or no difference (the null hypothesis), of obtaining a result equal to or more extreme than what was actually observed.
S. N. Goodman. Toward evidence-based medical statistics. 1: The P value fallacy. Annals of Internal Medicine, 130:995–1004, 1999.
edit: oh dear, and then the Ars article says "Individual experiments may be wrong five percent of the time," but that's exactly what p values do not measure. Statistics is hard.
For the purposes of understanding the definition, another way of looking at it is basically a 'statistical proof by contradiction':
1. Assume null hypothesis is true
2. Compute test statistic
3. Ask the question, "What is the probability of obtaining that test statistic or one more extreme?" (this probability is the p-value)
4. Pick a threshold (usually 0.05 but this is totally arbitrary)
5. If p < threshold then conclude the null hypothesis is false and reject it.
Reductio ad statistico absurdum.
Let's assume that 90% of results are false positives. Using a higher power experiment with even the same standard of p=0.05 would result in rejecting most of those, while hopefully keeping the majority of the true positives (due to the higher power of the new test). This would result in going from 90% false positives to less than half false positives, a considerable improvement.
It just seems that nobody ever advanced their career by trying to reproduce even a landmark result in their field.
This search for p=0.05 also leads to a lot of hair-splitting studies: take two diagnoses, say, usual and atypical ductal hyperplasia. Now, if you can find some constellation of parameters that define a middle category, say, "borderline ductal hyperplasia", you have a wide-open field to all sorts of p=0.05's, even if there's no change in treatment or outcome. You can say "cases previously characterized as UDH with a <parameter x> greater than <x> are 67% more likely to have <parameter y> (p=0.002)" because you lumped together a bunch of stuff that people already mostly agreed on anyway.
- Methodology - If you're mining for answers, and then presenting only the statistically significant, you will still find false positives on larger datasets, it will just take more work.
- Reproduction - If false positives were more fiercely chased down, then this would force researchers to be more careful. There is improvement along these lines.
Changing the p value is arbitrary, and may miss important results. I believe that encouraging reproduction of results, and reducing blind data mining is a better solution.
It does not take millions to check for basic mistakes. One person with computer and free afternoon is enough.
It is like heaving open-source, but without any source code.
This reminds me of the med researcher who rediscovered integration in 1993 and was cited many, many times:
Another step in the right direction is greater reproducibility, so that at least people can play with your analysis and see if what you did was the most direct and natural analysis, or if there are clear signs of playing with the parameters until you get the result you want.