It’s not just p=0.048 vs. p=0.052
statmodeling.stat.columbia.edu
statmodeling.stat.columbia.edu
More importantly, p=0.06 means that the researchers are honest. They could have easily p-hacked the results below 0.05 but chose not to. The opposite is true when p=0.049.
https://putanumonit.com/2018/09/07/the-scent-of-bad-psycholo...
> Here's the problem in a nutshell: If you run 1000 experiments over the course of your career, and you get a significant effect (p < .05) in 95 of those experiments, you might expect that 5% of these 95 significant effects would be false positives. However, as an example shown later in this blog will show, the actual false positive rate may be 47%.
> […] However, this is a statement about what happens when the null hypothesis is actually true. In real research, we don't know whether the null hypothesis is actually true. If we knew that, we wouldn't need any statistics! In real research, we have a p value, and we want to know whether we should accept or reject the null hypothesis. The probability of a false positive in that situation is not the same as the probability of a false positive when the null hypothesis is true. It can be way higher.
https://lucklab.ucdavis.edu/blog/2018/4/19/why-i-lost-faith-...
> Here's a more simple thought experiment that gets across the point of why p(null | significant effect) /= p(significant effect | null), and why p-values are flawed as stated in the post.
> Imagine a society where scientists are really, really bad at hypothesis generation. In fact, they're so bad that they only test null hypothesis that are true. So in this hypothetical society, the null hypothesis in any scientific experiment ever done is true. But statistically using a p value of 0.05, we'll still reject the null in 5% of experiments. And those experiments will then end up being published in scientific literature. But then this society's scientific literature now only contains false results - literally all published scientific results are false.
> Of course, in real life, we hope that our scientists have better intuition for what is in fact true - that is, we hope that the "prior" probability in Bayes' theorem, p(null), is not 1.
I think it'd be eye opening.
This is my opinion of course, it's not like I have any scientific basis for this.
I put a reminder to check back with you about this in a few months. Is Keybase your preferred contact method? I know it in name only, but I imagine I can figure it out.
You'd be given a prize/goal for publishable findings. Then you'd gradually introduce enough bias into the experiments to get something publishable, and then get hit with the reveal that "oh you were generating effects from random data, jerk".
(Okay, maybe that's overcomplicating it.)
"False positive" means that the effect or condition we're looking for is not true, but the experiment yields a true answer: the positive answer of the experiment is a falsehood. If the condition we're looking for is true, then there can't be a false positive. Even if the experiment yields a positive due to some flawed step, it's still a true positive.
This is just a language issue: a false positive of the rejection of the null hypothesis.
I'm all for criticism of p-values, but when I read a lot of critiques, I get to this point and simply stop reading.
No statistics text book that I've read assigns the magic value of p=0.05 and labels it as significant. All the ones I've read tell you to pick a p-value appropriate to your experiment. Yes, I get it that many social scientists don't have much of a clue and use 0.05 as some special threshold, but let's direct the criticism to the guilty parties, instead of blaming a statistical methodology.
I mean, we all know people who misuse the mean ("the average number of breasts a person has is 1") and ignore the shape of the distribution and the standard deviation. Yet we don't say "Let's stop using the mean!"
Those may be the hypotheses being tested, but they are not NULL hypotheses. NHST is not about testing the NULL hypothesis, it is about testing a non-Null hypothesis (the null hypothesis should always be incredibly boring and expected). There may be problems, but this blog post does not describe one.
Recently I worked with a client to interpret results from an A/B test where A performed better than B with 85% confidence (based on credible intervals, accounting for multiple comparisons). We therefore recommended A. In a group phone call, the client told her colleagues that our company doesn't know what we're talking about because 85% confidence of a difference isn't statistically significant (i.e. isn't 95% confident). We lost their business.
This was a shame because gathering the data for the experiment was expensive and the downside of making the wrong choice was low. It is often the case that taking on more risk makes more sense than hitting diminishing returns on shrinking p-values with extra sample.
The current trend of saying that "cutting p-values off at a specific value is bad" makes me worry. Now you can argue that your p=0.06 result shouldn't be rejected when really we should probably be pushing for stricter standards rather than inching towards looser ones. It also destroys the nice interpretation of p-values above. P-values were literally made to be cut off - if you want to stop doing that, you need to show me a coherent philosophy of what to do instead.
What I do think is true is the problem you have where part A of the experiment suggests X so you test X more directly in part B with a weaker but more specific test and get p=0.06 and now you can't publish. That's a dumb cutoff, clearly a p=0.06 test is likely to shift our belief towards X so it does nothing but bolster part A. Typically papers do this several times and the marginal 'failure' of one step should not sink the entire ship. This is a case where a Bayesian analysis might be more useful as it can incorporate weak evidence.
But the problem I see often is not that p-values are misused but that they were junk in the first place. For example, the widely-used DESeq2 (as well as some competitors in RNA-seq differential expression analysis) will happily spit out p-values of 10^-100 for an experiment with only four replicates in each of two conditions! There is no way you can get that level of evidence from just four replicates, even if the values are 0,0,0,0 and 1e6,1e6,1e6,1e6. The assumption of normality is reasonable near the mean but gets increasingly inaccurate in the tail, which is exactly where you end up when you do things like sort 30,000 tests by their p-values. In fact taking a p-value cutoff is probably the only reasonable thing to do here - that way you'll ignore the fact that it's absurdly small and just treat it as "small enough".
None of the great achievements of science (Newton, Darwin, Mendeleyev, etc.) were obtained on the basis of Popperian demarcationism/conjectures-and-refutations -- they were obtained by positing a large framework and patching together the empirical case for it.
Falsificationism isn't a stupid idea; it's even useful at a personal improvement level. But pharma or materials research use it because it tends to lead to good results, not because it's the very definition of what's worthwhile knowledge.
From the author of the article: "The general point reminds me of my dictum that statistical hypothesis testing works the opposite way that people think it does. The usual thinking is that if a hyp test rejects, you’ve learned something, but if the test does not reject, you can’t say anything. I’d say it’s the opposite: if the test rejects, you haven’t learned anything—after all, we know ahead of time that just about all null hypotheses of interest are false—but if the test doesn’t reject, you’ve learned the useful fact that you don’t have enough data in your analysis to distinguish from pure noise."
(https://statmodeling.stat.columbia.edu/2019/08/18/i-feel-lik...)
Your null should have a width. You should always be rejecting "effect is greater than some margin" which you should have to argue is greater than any bias you might expect in your experiment. There are always at least tiny biases.
The problem with a 0.048 and a 0.052 is not a mathematical one but an interpretation one. Reviewers are condition to be very skeptical of non-significant results and use “under power-ness” as a grounds for rejection. As a result, we get publication bias and p-hacking.
Machine Learning will (have) the same issue if the only thing that matters is hitting a certain level of accuracy given your model and data. This has been observed in Kaggle competitions over and over, you ask a group of people to find the best fit, and they'll, by learning your train, validation and test datasets.
As mentioned, problem is not p-value, or null hypothesis testing, the problem is journals who promoted the wrong incentive, and educators who were not aware of the consequences and propagated the wrong incentive (interpretation) to students.
Assuming the null hypothesis, that is precisely our expectation of finding something significant under the p < .05 rule. (That is, assuming that all papers try to falsely reject a true null hypothesis, then we expect 5% of the papers to be successful at that and get published.)
It wasn't until much later textbooks started to merge both. It may be worth to review Neyman and Pearson's attacks on Fisher in this matter.
I don't really follow this. Could someone clarify what is meant here? At what point would this author say something is not consistent with the null hypothesis?
Edit. Under Bayesian statistics testing the null hypothesis is a moot point as it becomes possible to directly model the distribution of the possible effects. Thinking of it as being able to look at a picture of something (the p-value) vs looking at a movie of it (the distribution of the effects).
Let's assume you've already decided in advance what "far" means.
Without moving either city from its current location, the same experiment can give you "very far" and "very close" in identical replications.
That is the frequentist point of view!
I'm completely lost here. How is 0.005 "dead center"? Are you assuming p = 0 is the center? Are there negative p-values I'm not seeing that somehow balance the positive ones?
How can a random variable that's strictly between 0 and 1 even follow a bell curve?
[0] http://www-ist.massey.ac.nz/dstirlin/CAST/CAST/HtestPValue/t...
Am I severely lacking sleep and going crazy or something? Maybe I should check back in like half a day to see what people have said, I feel like I must be completely confused right now because literally nothing I've read so far makes sense to me.
Here's an R example to play with:
pvals <- replicate(10000, {
x <- rnorm(100)
y <- rnorm(100)
t.test(x, y)$p.value
})
plot(density(pvals))
That will plot you a nice uniform line on [0, 1].(NB: I have no idea why OP talked about p values following a normal distribution. That doesn't make sense to me, and I think the post has been deleted.)
<< HypothesisTesting`
With[{n = 10000000, dist = NormalDistribution[]},
Histogram[Last[NormalPValue[RandomVariate[dist, n] - RandomVariate[dist, n]]], 500]]
Why is your x variable though? If H0 is true shouldn't your x be fixed?I think if you remove the subtraction though then you do get a uniform distribution -- in which case I see what the claim is, yeah. Wasn't really clear to me earlier but indeed, getting p = 5% means you have a 5% chance of getting observations that extreme, so I guess it is uniformly distributed!
(Disclaimer: I am not a real statistician....)
That's the very definition of a p-value! The mapping of data to p-values is chosen to have a uniform distribution of p-values when the data is distributed according to the null hypothesis. That's the property that makes p-values interesting.
> a P < 0.05 means that there is less than a 5% chance that the null hypothesis is true.
In other words, P(H0 | X) where H0 is the null hypothesis being true and X is the data observed. But that is not what a p-value is, they actually represent P(X | H0).
At which point, what is even the point of this statement?
The distribution will be different when then true hypothesis is not true (but you may also get non-significant results even if the null hypothesis is not true).
I’m not sure if that’s what you mean by “the actual distribution will be skewed in cases where it is statistically significant.”
For a identical normal populations, repeating an experiment produces p<=0.2 20% of the time, and produces p<=0.005 0.5% of the time.
A coin comes up on the same side 3 times in a row vs 8 times in a row...I have no idea why we should shrug and consider the plausibility of coin bias in these two cases about the same.
---
EDIT: I he means that in the case of an actual difference in the populations, p=0.2 and p=0.005 are both pretty likely outcomes.
When the populations are the same, p=0.2 and p=0.005 are quite different happenings.
This is because p-value methods doesn't worry very much about type II errors.
> P-values are shown to be extremely skewed and volatile, regardless of the sample size n, and vary greatly across repetitions of exactly same protocols under identical stochastic copies of the phenomenon; such volatility makes the minimum p value diverge significantly from the "true" one. Setting the power is shown to offer little remedy unless sample size is increased markedly or the p-value is lowered by at least one order of magnitude.
What you call "p-value" is a sample from the "p-value distribution" of your experiment.
Taleb shows you can sample a p-value of 0.05 when the actual "true" p-value is 0.12.
I do find it ironic though that this is so difficult to explain that I apparently have to read a paper to understand it... I would've thought the blog post was trying to explain things in simple terms...
IF the null is true, you're equally likely to get a p-value of 0.01 and 0.87.
And the argument will hold no matter what threshold you choose for rejecting the null hypothesis. You can choose to reject if p > X, and for any X, there will be values greater than X that, applying this meta-logic, are not statistically different from X.
> At what point would this author say something is not consistent with the null hypothesis?
Gelman's argument, I presume, is against the idea of significance testing as a whole. Declaring something "statistically significant" is in itself a very problematic thing, as it distills the entire phenomenon, the uncertainty surrounding the experiment, and the uncertainty surrounding the researcher's decisions to a single, binary conclusion.
Gelman is a Bayesian (perhaps the most famous modern Bayesian), and the Bayesian philosophy is to focus on producing a posterior distribution of the phenomenon being studied. I presume the alternative to significance and null hypothesis testing that he was suggest would be something closer to a model where people are reporting their priors/data/posteriors, and the discussion focuses around the implications and replication of those.
Isn't it though? The probability of this large (or larger) of a variance happening purely by chance[1]?
This article is highly critical, but the criticism goes over my head at least.
[1] assuming a normally distributed population
I'd guess that the original writer understands this, and that Gelman is only pointing it out because casual readers sometimes don't mentally retain the full baggage that the p-value carries.
For a popsci work, you could check out "how to lie with statistics", a classic.
Is it not? According to Wikipedia, it's "[...] the probability that, when the null hypothesis is true, the statistical summary [...] would be equal to, or more extreme than, the actual observed results." This sounds pretty much like "probability of happening by chance".
The difference is that, as highlighted in your quote, there is some null hypothesis that is assumed when discussing p-values.
For example: what is the probability of drawing x>2 when the underlying distribution is assumed to be a standard normal distribution N(0,1)?
The probability is small in this case, and could provide evidence to reject the null hypothesis (i.e. the distribution is not standard normal). It doesn't tell you about the probability of drawing x>2, it only gives evidence to reject (or not) the null hypothesis.
The wiki has more elaborate explanation. And probably better examples than mine.
It is the “[probability of happening] by chance” (as opposed to the [probability of not happening] by chance).