Not Even Scientists Can Easily Explain P-values
fivethirtyeight.com
fivethirtyeight.com
I do think that this book did an excellent job of explaining it:
What Is a P-Value Anyway? http://www.pearsonhighered.com/vickers/
Excerpt (starts at p.57 in the link above):
[I]f you do nothing else, please try to remember the following sentence: “the p-value is the probability that the data would be at least as extreme as those observed, if the null hypothesis were true.” Though I’d prefer that you also understood it—about which, teeth brushing.
I have three young children. In the evening, before we get to bedtime stories (bedtime stories being a nice way to end the day), we have to persuade them all to bathe, use the toilet, clean their teeth, change into pajamas, get their clothes ready for the next day and then actually get into bed (the persuading part being a nice way to go crazy). My five-year-old can often be found sitting on his bed, fully dressed, claiming to have clean teeth. The give-away is the bone dry toothbrush: he says that he has brushed his teeth, I tell him that he couldn’t have.
My reasoning here goes like this: the toothbrush is dry; it is unlikely that the toothbrush would be dry if my son had cleaned his teeth; therefore he hasn’t cleaned his teeth. Or using statistician-speak: here are the data (a dry toothbrush); here is a hypothesis (my son has cleaned his teeth); the data would be unusual if the hypothesis were true, therefore we should reject the hypothesis.
[...]
So here is what to parrot when we run into each other at a bar and I still haven’t managed to work out any new party tricks: “The p-value is the probability that the data would be at least as extreme as those observed, if the null hypothesis were true.” When I recover from shock, you can explain it to me in terms of a toothbrush (“The probability of the toothbrush being dry if you’ve just cleaned your teeth”).
I don't think this works so well because you are jumping from a qualitative statement (dry or not dry) to a quantitative ("at least as extreme").
It also isn't clear to me how you would apply your example given how most p-values are used. I read an article that says exercising makes you richer with a 95% confidence. So, as a layman, if I attempt to apply your example, it leads me directly into a common p-value mistake: the probability of me becoming richer if I exercise. 95%, right? Wrong!
That one quote wasn't really meant to take the layman from zero to "p-value master" in a handful of sentences.
I'd explain it like this:
It's easy to lie with statistics if you take them out of context. Maybe I have an idea that drinking flouridated water causes cancer, so I take a survey and find that 40% of people who drink flouridated water develop cancer. That number is meaningless without knowing how likely it is for people who don't drink it to develop cancer, which is about 40%, so the numbers match up 1:1, completely normal. In this case, 1 is the p-value, and it gives you a starting point for developing an experiment.
That is the frequency of cancer given not drinking fluoridated water. That is not a p-value.
Looking at a specific p-value and understanding its significance is only useful once they're actually looking at a study. But first they just want to know how it's derived and what it's good for.
1. p-values are intrinsically none-too-intuitive.
2. What p-values tell us isn't necessarily that interesting.
The problem is that the concept sounds more important than it is if you don't understand it. This is a bad failure state.
The issue is that the probability of some phenomena happening by chance is mostly orthogonal to the probability of that phenomena being true or whatever, but this seems outside the ability of most people to grasp.
P-values are quite intuitive. They are the ratio of observations made which agree with the null hypothesis to the total number of observations made.
Edit: If you are going to downmod, at least point out the error you perceive.
I think the resistance you're encountering stems from the fact that you're appealing to formality rather than intuition.
That depends on your hypothesis.
>The p-value is the probability of seeing a set of values
That is a circular definition. What is, "probability"?
P-values are not prescriptive statements about future observations, but are descriptions of actual observations.
>More specifically, the p-value is defined as the probability of...
Edit: A good exercise would be to start with the axioms[0] of probability theory and then derive the p-value for a simple experiment, using only those axioms (and your measured values).
Note that, probabilities of 0 or 1 are completely allowed under the theory and the computed value in your example is still consistent with my statement and probability theory.
An issue is that probability distributions (such as standard normal) don't exist in the axioms and so must be constructed from the axioms. So, a p-value test using a standard normal distribution is taking a shortcut that hides some of the calculations, making it appear that p-value is not actually counting a ratio, however, if you look at the full calculation from the axioms, p-value represents the ratio of agreeing observations to total observations. Basically, a single observation at the level of testing a standard normal distribution may actually represent multiple observations at the level of the math required to construct the full calculation (this entirely depends on the hypothesis being tested).
You are doing it wrong. A p-value based on a single observation is meaningless. P-values are much more useful when they are continuously computed on a running experiment with sequential observations.
See the work at the LHC for an example.
Σ P(E)
Each "unit" of the computation is weighted equally (though, the value of the units may differ).
You'll note that the notes don't derive the calculation of p-value. It merely gives a trivial example (curiously, the sum of two probabilities) with a fiat interpretation.
It's curious that my last post, merely quoting the third axiom and showing its properties plain, makes you think that I am deeply confused, when you cannot even refute trivial points and must resort to arguments (poor, in that they do not address our issue here) from authority.
> We will restrict ourselves to discrete sample spaces for now, to avoid some technical difficulties…
In the later lectures they take a look at sample spaces that are uncountably infinite. Good luck working with a sum there.
What reformulation of the probability axioms are you using to eliminate the sum in the third axiom?
The axioms may be trivially extended to bounded intervals (a la calculus).
Do you agree now that if the null hypothesis is true the p-value is uniformly distributed between 0 and 1? Or do you still think that "if that were true, p-value would be entirely useless"?
In any case, you were completely wrong about p-values a few months ago; maybe they are not so intuitive after all.
>Do you agree now that if the null hypothesis is true the p-value is uniformly distributed between 0 and 1?
If the null hypothesis is true, the p-value will converge to 1. (This makes complete intuitive sense as well, if you hypothesis is true, every observation you make should/will agree with it, making the ratio of agreeable_observations:total_observations = 1.)
You haven't shown anything where a true null hypothesis uniformly generates p-values between 0 and 1. Perhaps for a single observation you can only get a 0 or 1 value, and, perhaps for a over/under the average test (where every experiment will give 0 or 1) your sequence of p-values for each observation would be uniformly distributed, but p-values are generally computed as a normalized sum over many observations, in which case the value should converge to 1 if the null hypothesis is true.
One unintuitive aspect is that you might expect p-value to converge to 0 if the null hypothesis is false, however p-value is undefined when the null hypothesis is false.
Edit: Probably you don't care, but for the record:
"Since the value of x that defines the left tail or right tail event is a random variable, this makes the p-value a function of x and a random variable in itself defined uniformly over [0,1] interval, assuming x is continuous." [ https://en.wikipedia.org/wiki/P-value ]
"In statistics, when a p-value is used as a test statistic for a simple null hypothesis, and the distribution of the test statistic is continuous, then the p-value is uniformly distributed between 0 and 1 if the null hypothesis is true." [ https://en.wikipedia.org/wiki/Uniform_distribution_(continuo... ]
Most null hypotheses hypothesize a normally distributed measurement error. For any given distribution (continuous uniform), of course, the p-value means something different. That is what I have been arguing this whole time.
Here is a full derivation from a random blog, which I guess doesn't count either: https://shihho.wordpress.com/2012/11/27/pvalue_distribution/
Another one: https://joyeuserrance.wordpress.com/2011/04/22/proof-that-p-...
>https://shihho.wordpress.com/2012/11/27/pvalue_distribution/
This treatment commits the same error as the other poster (evanpw; I believe nonbel commits the same). The image at the bottom is only showing p-values for single observations. That is near useless (as the chart shows). An actual experiment requires repeated observation and analysis, not one-off observations.
npval <- pnorm( rnorm(nSim, mean, std.dev), mean, std.dev )
hist(npval, breaks=40, xlab=xlabel, main=title, col='gray')
R is probably the worst choice of language here, but my understanding of `pnorm` is that, when passed a list of numbers as the first argument (as `rnorm` returns), it will return a list of p-values (assuming normal distribution with given parameters), where the index in the list corresponds to the p-value, taking into account only that single observation (ie. it ignores all other observations in the list [ie. not how p-value hypothesis testing works {at least, in theory; real world practices may differ}]).Your second link appears to be the same argument (at least, they appear to use the same equations).
The charts are a nice visualization of how you can get a uniform distribution from a normal one (or vice versa) though.
* DESCRIPTION
*
* The main computation evaluates near-minimax approximations derived
* from those in "Rational Chebyshev approximations for the error
* function" by W. J. Cody, Math. Comp., 1969, 631-637. This
* transportable program uses rational functions that theoretically
* approximate the normal distribution function to at least 18
* significant decimal digits. The accuracy achieved depends on the
* arithmetic system, the compiler, the intrinsic functions, and
* proper selection of the machine-dependent constants.
*
So...yeah. I hope the constants the R authors picked work for your machine! const static double a[5] = {
2.2352520354606839287,
161.02823106855587881,
1067.6894854603709582,
18154.981253343561249,
0.065682337918207449113
};
const static double b[4] = {
47.20258190468824187,
976.09855173777669322,
10260.932208618978205,
45507.789335026729956
};
const static double c[9] = {
0.39894151208813466764,
8.8831497943883759412,
93.506656132177855979,
597.27027639480026226,
2494.5375852903726711,
6848.1904505362823326,
11602.651437647350124,
9842.7148383839780218,
1.0765576773720192317e-8
};
const static double d[8] = {
22.266688044328115691,
235.38790178262499861,
1519.377599407554805,
6485.558298266760755,
18615.571640885098091,
34900.952721145977266,
38912.003286093271411,
19685.429676859990727
};
const static double p[6] = {
0.21589853405795699,
0.1274011611602473639,
0.022235277870649807,
0.001421619193227893466,
2.9112874951168792e-5,
0.02307344176494017303
};
const static double q[5] = {
1.28426009614491121,
0.468238212480865118,
0.0659881378689285515,
0.00378239633202758244,
7.29751555083966205e-5
};
https://svn.r-project.org/R/trunk/src/nmath/pnorm.cEdit: Ah, a drive-by downmodder. How nice.
Try this in R. It calculates p-values for testing the hypothesis that two groups (n=10) come from the same distribution (normal with mean=0, sd=1). It does 100,000 such comparisons, then makes a histogram. This results in approximately a uniform(0,1) distribution of p-values:
hist(replicate(10^5,t.test(rnorm(10),rnorm(10),var.equal=T)$p.value))
If you increase to 10^6, 10^7, etc it will converge upon the uniform distribution. That is what people mean.
Edit:
Or something like this. Keep adding to the sample, the p-value will random walk between 0 and 1. The mean will converge onto 0.5:
n=10; Nsim=10000
a=rnorm(n); b=rnorm(n)
p=matrix(nrow=Nsim,ncol=2)
for(i in 1:Nsim){
p[i,1]=length(a)
p[i,2]=t.test(a,b,var.equal=T)$p.value
a=c(a,rnorm(n)); b=c(b,rnorm(n))
plot(p[,1],p[,2],ylim=c(0,1), ylab="pvalue", xlab="sample size")
}
High p-value is the first thing to look at a study. If it's high, then you can simply assume the findings are unreliable. But if it's low, it doesn't yet mean anything.
The p-values allow you to discard worst mumbojumbo quickly. That might not be interesting, but it's valuable. The thing to remember is that "we don't know" =/ "untrue".
For example, suppose the police are investigating the robbery of a home. They find a stranger's (call him person S) fingerprints all over the house. If S claimed he was innocent, a questioner might ask: "If you're really innocent, then why were your fingerprints all over the house?" This is another way of stating that P(Data | innocent) is small (similar to the p-value concept).
Since we often don't have a strong prior reason for believing a given individual is innocent or guilty, the p-value heuristic can be reasonable.
The formal connection to P(innocent | Data) is given by Bayes' theorem
That means that knowing the correct definition is no guarantee that you're actually using it correctly. (Odds are you aren't.)
In other words, I suspect the scientists that couldn't explain p-values in an intuitive way, didn't themselves fully grasp the subtleties of p-values.
That is what a p-value means, feasibility aside. There are also more intuitive ways to think about it, "false-positive rate" being one.
Suddenly, the abysmal state of nutritional epidemiology makes sense. How many labs do you suspect are running "red meat = death" studies?
Sadly, it only goes to point out the sad state of contemporary science. Especially in terms of nutrition, the only way to stay sane is to stick to the good old street-fighting Bayes - if popularly consumed product X was leading to cancer in any way relevant to your life, you'd see people dropping like flies all over the place.
I suspect you already know this, but it bears repetition, if only because it keeps us optimistic: science is a process. Even nutritional epidemiology will eventually find its way out of the abyss. :)
And yes: the ol' jellybean XKCD is quite a propos.
So as it stands now, I classify every new and contemporary (younger than 10 years) study in psychology, sociology and nutrition as "not science" and "quackery" by default, until proven otherwise. I also assume that any science news reporting in any field is an outright lie with an agenda. Rarely do I see a counterexample. Seriously, science reporting does way more harm than bad research itself.
I'd encourage you to be careful in dismissing the entirety of experimental psychology (full disclosure: I'm a cognitive neuroscientist, which is what a psychologist calls himself when he wants to distance himself from social psych).
The reproduction crisis doesn't touch all fields of psych in equivalent ways. When you reject the field as a whole, you're also rejecting stuff like this: http://cavlab.net/?lang=en
Hmm, p-values are data-dependent random variables. Do this simple experiment in R where the null hypothesis of no difference is as true as can be:
replicate(2,t.test(rnorm(10),rnorm(10))$p.value)
For example, I got:
[1] 0.8446991 0.3935238
So, the p-value can vary wildly (actually will follow a uniform(0,1) distribution) even if there were no real difference.
Experimenter A would tell you that "we would see a result like this 84% of the time anyway" and experimenter B would say 39%. If enough experiments are performed, eventually someone will calculate a very small p-value. So which experimenter should we listen to? Which experimenters should bother sharing their results? All are correct.
In other words, what should be done with that information about "even if there were no difference"? How should it affect our behavior, decisions, and beliefs?
https://en.wikipedia.org/wiki/Verification_and_validation
A p-value is a verification tool and nothing more, and yet scientists all too often take the "hunt for the p-value" as the only goal. Rather than fish through data for a p-value that ends up defining the scientific narrative in a paper, p-values need to be integrated into a much larger validation process driven by pre-defined scientific goals. It's no wonder that so many results are found to be specious given today's scientific culture.
If you only accept results with p-value <= 0.05, your false positive rate is at most 5%. Of course, if you had apriori knowledge that everything was a "null result" , then 100% of the results you accept would be false positives.
By Bayes' theorem, P(H|D) = P(D|H) * P(H) / P(D) . Therefore, given a fixed "prior probability" of a hypothesis P(H), a lower p-value (P(D|H)) implies a lower probability of the null being true. I think this relationship leads to a lot of confusion when thinking of p-values.
>"Ninety-seven percent of original studies had statistically significant results. Thirty-six percent of replications had statistically significant results; 47% of original effect sizes were in the 95% confidence interval of the replication effect size; 39% of effects were subjectively rated to have replicated the original result; [...] In cell biology, two industrial laboratories reported success replicating the original results of landmark findings in only 11 and 25% of the attempted cases" http://www.ncbi.nlm.nih.gov/pubmed/26315443
Edit:
For comparison, Dr. Oz was accused of fraud and called in front of congress for only being wrong half the time: https://www.washingtonpost.com/news/morning-mix/wp/2014/12/1...
Edit: This is just my experience at one university with 6-7 biological research labs. I regret using the generalization, apologies.
I many years of experience in helping people understand A/B testing. And I have found that no matter how many times you give a correct explanation, people will have this exact misunderstanding. Repeatedly.
So if you're A/B testing, here is something to consider. If you successfully set up a testing culture where you decide at a 95% confidence level, and are running one A/B test per week, you will make an average of 2-3 mistakes per year. In the long run, mistakes are unavoidable. You will make mistakes. The only question is how often you make mistakes, and how bad those mistakes are.
P-values manage to bound how often you make mistakes with the trick of most of the time saying you couldn't figure out an answer. THIS IS USELESS FOR A BUSINESS. A business needs to decide what text to use, even if they're not confident.
Your A/B tests should always produce a usable business answer. The question of interest is how often you make bad mistakes and how bad they are. And p-values are useless for that.
Not surprising. The correct interpretation of p-values was not discovered until ~2 years ago. If you search around the internet you will find out that, according to the discoverer, he couldn't find a stats journal to publish it. The number of people who know the correct definition must be vanishingly small.
Thank god for arxiv, or I would have never understood p-values: http://arxiv.org/abs/1311.0081
tldr: Calculating the p-value + sample size is a lossless compression algorithm. The thing being compressed is a likelihood function (way of describing effect size).
Here's an intuitive explanation: a p-value is the probability of getting your experimental results given that your hypothesis is wrong.
The p-value won't tell you if the coin is fair or not, but it can tell you the probability that the coin is fair.
I think that's sort of the point that the article is making actually. That high probability does not imply truth. There are other non probabilistic ways to verify that coin is unfair, for example by looking at the density throughout the coin.
Even if you look at the density of the metal throughout the coin, there's still a chance I've altered your device to report the coin is fair. Or a passing microsingularity decided to play games with the scanning beam. Or you're just imagining the whole thing.
That's not to say one should despair that the world is unknowable. One only has to get used to the fact that, in practice, "true" just means "extremely, extremely likely".
But yeah truth is tricky.
Suppose I have a coin and flip HHHHH. Can you tell me the probability the coin is fair? No, it's fundamentally unknowable. We can say that a fair coin would have a 3% chance of flipping HHHHH (the p-value), but we can't say with what probability our coin is fair.
That's a common misconception. Actually it's the probability of getting your experimental results given that your null hypothesis is right.
Consider an experiment where your hypothesis is that cold temperatures cause the common cold. This is a good example for a thought experiment because we "know the answer" in a way (there have been a lot of experiments on this). The null hypothesis in this case is that cold temperatures are uncorrelated with incidence of the common cold.
You place people in isolation in cold areas and a control group in warm areas, and study how many get the common cold. None of the people who didn't already have colds get the common cold: because they are in isolation and the common cold is caused by rhinoviruses (which they can't get because they are in isolation), you get exactly the same results.
This disproves the hypothesis, but it does not prove the null hypothesis, that cold temperatures and the common cold are unrelated. Cold temperatures are, in fact, related to the common cold.
Try a second experiment: you place people in groups of five in cold areas and in warm areas, and discover a moderately high correlation between cold temperature and incidence of the common cold. This disproves the null hypothesis. But the simple hypothesis that cold temperature causes the common cold has also been disproven by your first experiment.
The reason for this is that the correlation between cold temperature and the common cold is a dependent correlation: given that rhinovirus is present in the system cold temperature is correlated with incidence of common cold (rhinoviruses reproduce ideally at temperatures significantly lower than human homeostatic temperature).
The null hypothesis is not just a statement that there is no independent correlation, it's a statement that there is no independent or dependent correlation. As such, the null hypothesis is an extremely broad hypothesis which is impossible to practically prove. This is why there's such a focus on finding correlations rather than finding non-correlations: you aren't going to prove the null hypothesis.
Then research with high p-value would just mean "we tried, and we still know practically nothing".
Or, suppose we have some data and want to use it to estimate the value of some number b. It can be that if we can assume something about b, e.g., that b = 0, then we can calculate the probability distribution of our estimate of b. This assumption about b is the null hypothesis. Intuitively it is a hypothesis that there is no effect, that what we thought might have happened didn't, was a null effect.
We can make two mistakes. We can reject the null hypothesis when it is true -- this is called Type I error. Or we can accept the null hypothesis when it is false -- this is called Type II error.
With our null hypothesis, we get and look at the distribution of our estimate of b and see where our actual estimate is in that distribution. The p-value is the probability, from the distribution of our estimate, of getting an estimate as far or farther from our null hypothesis value for b as we did. So, if the p-value is really small, say, 1%, then we can reject the null hypothesis, that is, say that it is false, and be wrong only 1% of the time.
E.g., if our null hypothesis is that b = 0 and our estimate of b is 10 and from the distribution of our estimate a value of our estimate being as far as 10 from b = 0 is 1%, then we reject that b = 0 and conclude that it b is not zero and are wrong only 1% of the time. If the probability of our estimate being greater than or equal to 10 is 1% and we reject the null hypothesis, then we conclude that b > 0.
This is all just hypothesis testing in statistics 101. Will also want to know about the power of a test, the t-test, the F ratio, the chi-squared test, and resampling and distribution-free tests.
There is more detail in
No, see here: "Of critical importance, as Goodman (1993) has pointed out, is the extensive failure to recognize the incompatibility of Fisher’s evidential p value with the Type I error rate, α, of Neyman–Pearson statistical orthodoxy."
P Values are not Error Probabilities http://www.uv.es/sestio/TechRep/tr14-03.pdf
I never saw a clean definition of p-value or alpha -- probability of Type I error.
In my study of statistics, p-value and alpha are essentially the same. Maybe a difference: A p-value is what a researcher has in mind before the hypothesis test as a 'cutoff value', say, 5%, for alpha while alpha is whatever the statistical hypothesis test says, for the data and the test, is the probability of Type I error, e.g., 3.141592653 or some such random number. So, since the 3.14 is less than the p-value of 5%, the experimenter rejects the null hypothsis. MAYBE this is the difference, if any, between p-value and alpha, in which case we're talking a triviality.
The paper you referenced also starts to get wound up over beta = 1 - alpha as the power of a test. Okay, but that little equation is the definition of power. So,if have some data and a p-value or alpha in mind, and have several candidate statistical hypothesis tests, say, some parametric and the others distribution-free, etc., then will want to use the test with the smallest probability of Type II error, that is, the largest beta. Okay. Now are done with beta, and no more strain or struggle needed.
A biggie point is the real role of the null hypothesis -- it lets us calculate some probabilities that, otherwise, we would not have assumptions enough to do.
If there was a fight between Fisher and Neyman-Pearson nearly 100 years ago, then I'm sorry, but now I have no sympathy for whatever the heck, if anything, they were arguing about then.
On re-reading this I agree. They do not make a clean comparison. They do define them though:
p-value: "the probability of the observed and more extreme data given the null"
alpha: "α is the long-run frequency of Type I errors"
To begin with, you appear to have these reversed. Second, the p-value is dependent on the data, it will be different from experiment to experiment. The exact same experiment can give you p=0.001 one time and p=0.32 the next time. Which is the probability of type I error?
This conditional probability is conditioned on the observations we used; since those are random variables, so is the conditional probability. Then the 0.32 and 0.001 are values of this random variable on the two trials. The expectation of this random variable will be the probability of Type I error alpha.
I strained and tried to find a difference between p value and alpha, but in current usage I see no difference. My guess is that people prefer p-value because it abbreviates probability while alpha does not.
Maybe some people want to say that
alpha = E[E[p value|given data]]
but this is not very operational. Maybe it is what people mean in which case, right, alpha is the expected value of p-value.
So, for a given experiment and hypothesis test, can draw a graph with alpha on the X-axis and beta (probability of Type II error) on the Y-axis. Then the graph shows beta as a function of alpha. So, typically the graph runs from (0,1) to (1,0) and is convex. A more powerful test is lower at each value of alpha. A perfect test is just the point at the origin. A trivial test is the one get by ignoring the real data and using a random number generator, and here the curve is just the straight line from (0,1) to (1,0) and is useful for some cases of interpolating when the available values of alpha are discrete, e.g., in resampling plans for distribution-free statistics.
Right, beta and power = 1 - beta don't have much to do with alpha or p-value except for the point that for a given statistical test in a specific context, e.g., distribution of the data and number of data points, there is that curve I described that relate alpha and beta, and, thus, also 1 - beta, exactly. As I outlined, a different statistical test, maybe parametric instead of distribution-free, can have a different curve and more power for a given alpha.
From the paper: alpha= probability of type I error
So, there are two equations:
p value= p(alpha|data)
alpha = E[E[p value|given data]]
Substituting we get:
alpha = E[E[(alpha|data)|given data]]
What is the difference between "given data" and "data"? How do you isolate alpha?
alpha = E[E[P(alpha|data)|given data]]
Now we can simplify. Say x=alpha, y=data, z="given data", the E operator can be an arbitrary function f, and applying function f twice is denoted as f#. We get:
x=f#(P(x|y)|z))
We know P(a|b)=P(a & b)/P(b) and p(a & b)=P(a)P(b), so
x=f#{[P(z)P(x)P(y)]/P(y)}=f#{P(z)P(x)}
Denote the inverse of f# as F, then:
F(x)=P(z)P(x)
Now, z was "given data" which has probability 1. So,
F(x)=P(x)
In other words, there is method to your madness.
E[X|Y]
is the almost sure unique random variable that integrates like X over all events in the sigma algebra generated by Y.
In particular
E[X] = E[E[X|Y]]
Details follow from the Radon-Nikodym theorem in, say, Rudin, Real and Complex Analysis where Rudin give the von Neumann proof. Details are also in the standard texts on graduate probability by Loeve, Neveu, Breiman, Chung, and a few more.
Well, you were correct about this above. You may be onto something here in terms of "actual use".
The probability of type I error of what exactly? Could you be more explicit?
For example: we want to perform a test at the 5% significance level. The procedure is simple: we reject the null hypothesis when p<alpha=0.05. If the null hypothesis is true, in 1000 trials on average:
- we get 0<p<0.01 10 times, we reject the null hypothesis
- we get 0.01<p<0.02 10 times, we reject the null hypothesis
- we get 0.02<p<0.03 10 times, we reject the null hypothesis
- we get 0.03<p<0.04 10 times, we reject the null hypothesis
- we get 0.04<p<0.05 10 times, we reject the null hypothesis
- we get 0.05<p<0.06 10 times, we cannot reject the null hypothesis
- we get 0.06<p<0.07 10 times, we cannot reject the null hypothesis, etc.
The type I error of the procedure is (by construction) 5%: when the null hypothesis is true, we have a rejection (false positive) on average 50 times in 1000 trials.
The "type I error rate" is a property of the test, not a property of the outcome. Using your example, we perform the test once and:
- we get p=0.001, we reject the null hypothesis. The Type I error rate of the test is 5%
- we get p=0.32, we do not reject the null hypothesis. The Type I error rate of the test is 5%
Now, could you please fill in the dotted lines?
- we get p=0.001, we [ reject the null hypothesis / do something else]. The Type I error rate of .................. is 0.1%
- we get p=0.32, we [ do not reject the null hypothesis / do something else ]. The Type I error rate of .................. is 32%
This is completely wrong: you're confusing the alpha and the p-value. Maybe you meant that alpha (5% in your example) is the preset type I error rate (significance level) and the p-value (here 3.14%) is calculated from the data. But conceptually the alpha and the p-value are different things and shouldn't be mixed. 3.14% is definitely not "for the data and the test, [...] the probability of Type I error".
If you're doing hypothesis testing, you calculate 3.14 which is less than 5 (alpha) and you reject the null hypothesis. If the calculation had yielded 0.69 instead of 3.14 you would also reject the null hypothesis. But you cannot consider the second result to be "stronger" than the first one. They are both below the threshold and that's all that matters.
If you're using p-values as evidential measures, the second result is indeed "stronger" because 0.69<3.14. But in this framework there are no error rates involved.
> The paper you referenced also starts to get wound up over beta = 1 - alpha as the power of a test.
I've not looked at the paper carefully, but it's unlikely. Power = 1 - beta, it is not directly related to alpha. (Maybe it's just a typing error, I'm not sure what is your point. I mention it mostly for the benefit of other readers that might be confused by your definition.)
The motivation for the p-value is to tell you what the probability of a fluke chance this would be.
The way to do better is to run the experiment over and over, flipping this coin 1000 times again, and again. However, since science experiments are extremely expensive, instead we only run them one time, and try to 'make sure this wasnt some fluke chance'.
http://www.sciencebasedmedicine.org/psychology-journal-bans-...
In the context of a regression analysis, I said, "P-values indicate the chance that the apparent effect of the variable is from random fluctuations in the data instead of the variable itself."
It's impossible to know the odds of an apparent effect coming from random data fluctuations or a real effect. The p-value cannot tell you these odds. The p-value only tells you one half of that---the odds that random fluctuations would cause an apparent effect.
"Some scientists can not easily explain p-values" should be a more accurate, and much better, title.
Does this make sense? If the p-values are not very good at conveying "confidence in the prediction," is this another argument in favor of a more Bayesian approach to statistics? Any thoughts would be appreciated!
From the first sentence of the parent comment: "From what I understand, p-values are typically used by frequentist method for verifying that the "prediction" seems correct."
If my reply seemed glib, it was because I thought the question was so far off base, that graffitici almost certainly hadn't read the article. To relate to the topic under discussion, the distribution of possible comments produced by people who have carefully read the article is extremely unlikely to result in a post as extremely misinformed as the one above.
If there was a simple statistical misunderstanding, perhaps that would be worth clearing up. Instead, the parent comment used some statistical words, but was very nearly incoherent. Perhaps I shouldn't have said anything if I wasn't willing to write a protracted essay about the differences between frequentist and Bayesian inference.
I disagree. There's an alternative hypothesis that could explain a comment like that - the author harbors some confused views about frequentist and Bayesian approaches to statistics. I'm inclined to believe that this hypothesis is right, because I recognize that comment as something I'd write myself back when I was more confused about this topic. Reading the article is unlikely to affect this particular issue.
> Perhaps I shouldn't have said anything if I wasn't willing to write a protracted essay about the differences between frequentist and Bayesian inference.
I think even one sentence explaining the gist of the author's confusion would be enough. Plenty of essays have been written on the topic, but one needs to be pointed in their general direction in order to benefit.
On the second point, frequentists very commonly split training and testing datasets, and so this practice is pretty much orthogonal to whether you're using Bayesian or frequentist methods.
The Bayesian approach would be to say that the thing you generally want to know isn't {the probability of getting this result by chance if the hypothesis is false}, it's the probability the hypothesis is true, given everything you know up to and including the new data.
So they would calculate p(hypothesis is true given the new data) using Bayes theorem -- which requires inputting what they thought p(hypothesis is true) was at the start of the experiment, before the new data came in.
http://www.yudkowsky.net/rational/bayes is a decent explanation.
Bayesian stats aren't related to training and test sets, like in machine learning; scientific experiments rarely partition data like that. A Bayesian analysis answers the question, "Given the evidence I just collected, how should I adjust my estimate of how likely something is?" But Bayesian analyses have issues too: the biggest is not knowing what your initial, "prior" probability should be. (After all, that's usually what you're trying to find out!) With something simple like a coin flip, you have a strong prior assumption of 50% probability of heads.
And neither Bayesian nor frequentist methods address effect size. If you collect enough data, you can detect extremely small (real) effect sizes that pass statistical tests, but are still meaningless outside the context of publishing a paper. Rather than analysis type, we should incorporate more discussion of effect size, especially as the number of data points increase.