P-Values are not Error Probabilities (2003) [pdf]
uv.es
uv.es
Case in point #1: debating whether a P value of 0.051 versus 0.049 is a major difference in the significance test.
Case in point #2: a P value of <0.001 but with extremely low differences in means between populations. With enough data, everything is significant!
End rant. :)
This is actually a terrible problem. Most people, I think, do this because of nervous cluelessness: they are sure there is some theory that says what it means, but they aren't sure they understand it, so they insist on following some strict rules, whether they make sense or not.
I think the idea is that if you follow a strict rule given in a textbook (even if wrong or inapplicable), that makes people feel safer than if they try to figure the rule out for themselves. At least if they follow a rule, they aren't personally responsible for the outcome.
It can sometimes be quite difficult to explain why a certain rule should not be applied, because the same reasons for why they didn't understand the rule in the first place also make it difficult to understand explanations given by someone else.
I took my bachelor in business and we were always told to make rational business decision. That is, informed decisions, based on data.
As for actual statistics, we were told 2 rules of thumbs: You need at least 30 people in a group before you can make meaningful statistics and it is significant if P < 0.005
Your statement contradicts the introductory paragraph of the article:
> "...the outcomes of these tests are mistakenly believed to yield the following information:... the probability that an initial finding will replicate; ..."
The p-value is the likelihood that the observed effect is due only to chance, not a measure of the repeatability of a result.
I much prefer how machine learning folks tend to approach predictive accuracy, though I guess that's not quite the same as understanding relationships between specific variables while controlling for others.
What's going on in the academia is a classical example of perverse incentives. If your career depends on churning out soul-crushingly bad or insignificant papers that fulfill some superficial criteria of "scienceness" then that's what you'll do. It's dangerously close to cargo cult science.
PDF link: http://www.stat.columbia.edu/~gelman/research/published/sign...
"I use pictures from the ESCI software to give a brief, easy account of the Dance of the p Values. The simulation illustrates how enormously and disastrously variable the p value is, simply because of sampling variability. Never trust a p value!"
1. At least at the beginning, it focuses excessively on the historical aspects of statistics. For example, it says that "most applied researchers are unmindful of the historical development of methods of statistical inference, and of the conflation of Fisherian a nd Neyman–Pearson ideas." To me, statisticians shouldn't /have/ to understand the history at all. For example, as a physicist, there is absolutely no need for me to understand the evolution of Ampère' theories, Faraday's theories, Maxwell theories, etc. to apply the laws of electricity and magnetism correctly.
2. The difference between p and alpha is central to the paper, but it doesn't seem to have a cogent explanation of what that difference is. (It's very clear who advocated for one and who advocated for the other, but that's not why the difference is important.)
Here's the idea. If you set alpha = 0.05, you will declare statistically significant any result that gets a p value of 0.05 or less. When there is no statistically significant difference to be found, you will have a 5% chance of falsely detecting one.
But crucially, this applies on average to all tests you conduct with this alpha level. Even if an individual test gets p = 0.000001 or p = 0.04, the overall false positive rate will be 5%.
More succinctly, it doesn't make sense to ask for the false positive rate of a single test. What does that even mean? You can only ask for the false positive rate of a procedure you use many times. So you can't get p = 0.01 and declare this means you have a false positive rate of 1%.
2. As far as what the difference is, I have to admit I have not found a memorable phrase to explain it. Both Goodman in "Toward Evidence-Based Medical Statistics" and this paper give a similar wording of the difference, which i would rephrase as follows.
The p-value works by inference, taking the data and assigning a probability to that data, not to an hypothesis. It is to be used to corroborate our disbelief in the null, i.e. informally. To use inference to test hypothesis one must use Bayes' Theorem, i.e. Bayes factors, and introduce prior probabilities. Hypothesis testing à la Neyman and Pearson is deductive process, in which one does assign probabilities to the hypotheses, paying the price of only being able to minimize the errors commited, not to draw inferences.
One should also confront "In Fisher’s approach the researcher sets up a null hypothesis that a sample comes from a hypothetical infinite population with a known sampling distribution." with "Neyman–Pearson results are predicated on the assumption of repeated random sampling from a defined population."
Given the natural predisposition by students to works in an inferential manner, it may be wise to bite the bullet and teach Bayes factor instead of the p-value cargo cult (this being a pedagogical choice, not an assessment of frequentism vs. bayesianism or a critique of p-values as conceived by Fischer).
"The Cult Of Statistical Significance":
http://www.amazon.com/Cult-Statistical-Significance-Economic...
It basically goes through a bunch of examples, mostly in Economics, but in medicine also (Vioxx) where statistical significance has failed us and people have died for it. As someone who works with statistics for a living, I found to book interesting - but it was pretty depressing to find out that most scientist are using t-test and p-values because it seems to be the status quo and it is the easiest way to get published. The authors of this book suggest a few different things -- publishing the size of your coefficients and using a loss function. In the end, they make the point that statistical significance is different than economic significance, political significance, etc.
http://myweb.brooklyn.liu.edu/cortiz/PDF%20Files/Misinterpre...
TLDR: 80% of methodology instructors have a misconception about significance. Scientific psychologists and students perform even worse.
"p’s and α’s are not the same thing; they measure different concepts"