I don't really follow this. Could someone clarify what is meant here? At what point would this author say something is not consistent with the null hypothesis?
I don't really follow this. Could someone clarify what is meant here? At what point would this author say something is not consistent with the null hypothesis?
And the argument will hold no matter what threshold you choose for rejecting the null hypothesis. You can choose to reject if p > X, and for any X, there will be values greater than X that, applying this meta-logic, are not statistically different from X.
> At what point would this author say something is not consistent with the null hypothesis?
Gelman's argument, I presume, is against the idea of significance testing as a whole. Declaring something "statistically significant" is in itself a very problematic thing, as it distills the entire phenomenon, the uncertainty surrounding the experiment, and the uncertainty surrounding the researcher's decisions to a single, binary conclusion.
Gelman is a Bayesian (perhaps the most famous modern Bayesian), and the Bayesian philosophy is to focus on producing a posterior distribution of the phenomenon being studied. I presume the alternative to significance and null hypothesis testing that he was suggest would be something closer to a model where people are reporting their priors/data/posteriors, and the discussion focuses around the implications and replication of those.
> P-values are shown to be extremely skewed and volatile, regardless of the sample size n, and vary greatly across repetitions of exactly same protocols under identical stochastic copies of the phenomenon; such volatility makes the minimum p value diverge significantly from the "true" one. Setting the power is shown to offer little remedy unless sample size is increased markedly or the p-value is lowered by at least one order of magnitude.
What you call "p-value" is a sample from the "p-value distribution" of your experiment.
Taleb shows you can sample a p-value of 0.05 when the actual "true" p-value is 0.12.
I do find it ironic though that this is so difficult to explain that I apparently have to read a paper to understand it... I would've thought the blog post was trying to explain things in simple terms...
Edit. Under Bayesian statistics testing the null hypothesis is a moot point as it becomes possible to directly model the distribution of the possible effects. Thinking of it as being able to look at a picture of something (the p-value) vs looking at a movie of it (the distribution of the effects).
Let's assume you've already decided in advance what "far" means.
Without moving either city from its current location, the same experiment can give you "very far" and "very close" in identical replications.
That is the frequentist point of view!
For a identical normal populations, repeating an experiment produces p<=0.2 20% of the time, and produces p<=0.005 0.5% of the time.
A coin comes up on the same side 3 times in a row vs 8 times in a row...I have no idea why we should shrug and consider the plausibility of coin bias in these two cases about the same.
---
EDIT: I he means that in the case of an actual difference in the populations, p=0.2 and p=0.005 are both pretty likely outcomes.
When the populations are the same, p=0.2 and p=0.005 are quite different happenings.
This is because p-value methods doesn't worry very much about type II errors.
IF the null is true, you're equally likely to get a p-value of 0.01 and 0.87.
At which point, what is even the point of this statement?
The distribution will be different when then true hypothesis is not true (but you may also get non-significant results even if the null hypothesis is not true).
I’m not sure if that’s what you mean by “the actual distribution will be skewed in cases where it is statistically significant.”
I'm completely lost here. How is 0.005 "dead center"? Are you assuming p = 0 is the center? Are there negative p-values I'm not seeing that somehow balance the positive ones?
How can a random variable that's strictly between 0 and 1 even follow a bell curve?
[0] http://www-ist.massey.ac.nz/dstirlin/CAST/CAST/HtestPValue/t...
Am I severely lacking sleep and going crazy or something? Maybe I should check back in like half a day to see what people have said, I feel like I must be completely confused right now because literally nothing I've read so far makes sense to me.
Here's an R example to play with:
pvals <- replicate(10000, {
x <- rnorm(100)
y <- rnorm(100)
t.test(x, y)$p.value
})
plot(density(pvals))
That will plot you a nice uniform line on [0, 1].(NB: I have no idea why OP talked about p values following a normal distribution. That doesn't make sense to me, and I think the post has been deleted.)
<< HypothesisTesting`
With[{n = 10000000, dist = NormalDistribution[]},
Histogram[Last[NormalPValue[RandomVariate[dist, n] - RandomVariate[dist, n]]], 500]]
Why is your x variable though? If H0 is true shouldn't your x be fixed?I think if you remove the subtraction though then you do get a uniform distribution -- in which case I see what the claim is, yeah. Wasn't really clear to me earlier but indeed, getting p = 5% means you have a 5% chance of getting observations that extreme, so I guess it is uniformly distributed!
(Disclaimer: I am not a real statistician....)
That's the very definition of a p-value! The mapping of data to p-values is chosen to have a uniform distribution of p-values when the data is distributed according to the null hypothesis. That's the property that makes p-values interesting.
> a P < 0.05 means that there is less than a 5% chance that the null hypothesis is true.
In other words, P(H0 | X) where H0 is the null hypothesis being true and X is the data observed. But that is not what a p-value is, they actually represent P(X | H0).