You assume a null hypothesis that (usually) represents the status quo of no influence between the theory and the data. You then collect data. The p value then describes the probability of that data aligning with / being as a result of the null hypothesis. In other words, a p <= 0.05 says that you have <5% chance that the data came from the theory stated in the null hypothesis - that is, you have a 95% confidence that you can reject the null hypothesis in favour of your new theory.
Is that correct? I may have minced terms there because my stats training is woefully inadequate, but I think I adequately conveyed the concept?
Computing the probability that the data came from the theory stated in the null hypothesis would require a (Baysian) prior.
Also, Tloewald's reply is completely and inexorably wrong. Tloewald seems to want a Bayesian answer, which frequentist statistics can't give you.
"Given the evidence, there is a >=5% probability of the null hypothesis being true"
and
"There is a >=5% probability that if the null hypothesis were true, that your data would be at least as extreme"
The only difference I see is how you avoided saying anything about the null hypothesis, but I don't see how you can avoid saying anything about it.
if the h0 were true, then the probability of the result is unlikely, how can you not conclude that h0 is unlikely? What step are you missing other than collecting a preponderance of evidence against it?
The article never enters into this distinction. It makes it clear that people misinterpret evidence against the null hypothesis as evidence of the alternative, which is a false dichotomy.
I am confused. I also have sympathy for Tloewald at this point.
The two quantities are related to each other via Bayes rule:
P(H0|D)=P(D|H0)P(H0)/P(D)
So indeed, as P(D|H0) goes down, so does P(H0|D). But if P(H0)/P(D) is sufficiently large, you can easily have P(H0|D) high while P(D|H0) is low.
I too have sympathy for everyone confused by frequentist stats - they tend to answer the exact opposite question that one really wants answered. In contrast, Bayesian stats tend to answer the question that most people ask.
What does P(D) mean?
I read that is, the probability of Data being true.
edit: to clear up my meaning.
I mean, it makes sense to me to ask "What is the probability of getting this data, given that the null hypothesis is true"
and "what is the probability of the null hypothesis being true, given this data"
but I don't know how "this data" evaluates on its own. I can't picture that
does it mean, how authoritative is the data? Maybe that's it.
edit: OK never mind I kinda worked it out on my own.
P(D) = P(D|H0)P(H0) + P(D|H1)P(H1)Lastly, if you're in a situation where you're performing a large number of tests, always correct for multiple testing (i.e. false discovery rate or some similar method). Also, you can take advantage of your large data set to construct a negative control (e.g. by shuffling samples, the exact method will vary) and verify that your chosen statistical test gives a flat distribution of p-values, which is the expected result when all of the null hypotheses are true. If you have an excess of small p-values in the null data set, this indicates your test is producing false positives and is not reliable (presumably because one or more assumptions have been violated).
Often the hypothesis testing framework is stated something like:
H0: µ = 0 (null hypothesis)
Ha: µ ≠ 0 (alternative hypothesis)
When you reject H0, it means that you can be somewhat confident that there was was some kind of distortion in your data that moved the mean away from (in this case) 0.You can create a number of theories that purport to explain this mechanistically, but you'll often need particular setups like a randomized controlled trial that can eliminate alternative explanations. When you've eliminated all the competing reasonable hypotheses, there can only be one. If you can use that hypothesis to make non-obvious predictions, that's further proof that it's right as well.
Hypothesis testing is there to tell you when to take an effect seriously, it doesn't tell you whether your explanation is right outside of very carefully constructed circumstances (i.e. where if you see a particular effect, only one theory can explain it).
For practical purposes NHST is a function that returns either H0 or Ha.
Say you have a confidence interval of 95% or higher (p value <= 5%), then the right thing to say is:
95% of the variance within the data is explained by the model (you've come up with).
I think it is ridiculous that the main focus always seems to be arcane properties of the statistical algorithm and not the answer that it delivers.
I said the theory -- by which I meant the non-null hypothesis -- is not endorsed by a low p value (indeed, this is a major point of the article). A low p-value says "hey they doesn't look like random data" not "your brilliant hypothesis is probably true". The data might not look random because of a methodological error, outright fraud, or a confound.
This is particularly important when you consider people looking at data over and over again trying to find "an effect". Theoretically, the tests are supposed to get tougher and tougher each time you examine the data (add one degree of freedom) but in practice this doesn't happen. It doesn't matter much with large data sets, but the social sciences often use datasets where n is roughly 100, and you might only have 20 subjects in a cell.
The same is true in Bayesian statistics, and even simple formal reasoning with no statistics in sight. If you make wrong assumptions, you'll get the wrong result.
The only thing you can expect statistics to do is help you change your opinion about the relative merits of opposing theories. If both your opposing theories are wrong, you will still be equally wrong.
The true flaw with frequentist statistics is that it goes out of it's way to hide this fact from you. In contrast, Bayesian stats forces you to explicitly choose a prior, enumerate your assumptions, and accept that your conclusion is based on these things.
"the problem is that, for example, a 95% confidence interval does not indicate that the parameter of interest has a 95% probability of being within the interval. Rather, it means merely that if an infinite number of samples were taken and confidence intervals computed, 95% of the confidence intervals would capture the population parameter"
And I say, "What? If out of every 100 random samples, in 95 of them the parameter is in the interval, then surely the probability of the parameter being in the interval is 95% by definition?"
Where was the catch? I remember there was one (which is enough for practical purposes because it means I won't say in a paper that the probability of the parameter being in the interval is 5%) but I feel dumb for not being able to see at first glance something that is supposed to be basic statistics...
If I have a variable that is always positive, then I could have a weird procedure to generate confidence intervals that gives me the interval [-inf,0] 5% of the time and the interval [0,inf] 95% of the time.
This would meet the definition of confidence intervals perfectly, and yet when I get [-inf,0] the real probability of the parameter being in the interval is 0%, and when I get [0,inf], it's 100%.
I wonder how large this discrepancy may be in practice (as this is obviously a made-up extreme case).
As a completely fabricated example, suppose the true proportion in a coin flip experiment is 40%. If my confidence interval is [.45, .65], what's the probability that .40 is in [.45, .65]? It's 1. The probability that a fixed, but unknown parameter will lie in any confidence interval will be either 0 (it isn't in the interval) or 1 (it is). The _proportion_ of times the interval contains the true parameter is the level of confidence (95%).
To your always-positive example, that procedure is not particularly weird. There's always a balancing act with CI's about length and confidence level (otherwise, I could choose all reals as my interval and get 100% confidence level). That your [0, inf) interval has 100% coverage means that you could probably shrink that interval so that it has finite upper bound without losing more than 5% confidence. Hard to say without a specific distribution in mind or mild assumptions, but an application of either Markov's or Chebyshev's Inequality would allow you to make really loose bounds with only relatively minor assumptions.
On the other hand credible intervals express `P(X ∈ [A, B] | A=a,B=b) = 0.95` (or more generally `P(X ∈ [a,b] | the data) = 0.95`). The latter is what is intuitively meant by "95% probability" of the true parameter being in the interval, because you do know a and b but not the parameter.
The example with random sampling of confidence intervals from {ℝ⁺, ℝ⁻} is indeed a good illustration of the difference.
In general, frequentists just don't like to talk about the properties of this particular sample, only about long-term frequencies – hence the name. Why? Because they object to the idea of probability as a degree of belief rather than as an objective measure, and given that attitude the statement that "there's a 95% probability the parameter is in this interval" doesn't make any sense: either it's in the interval or it isn't.
Huh? Unless you can bound the set of potential results, this isn't possible. Say I want to estimate the half-life of some material (bounded below, but not above). A uniform prior doesn't exist. How will the credible interval relate to the confidence interval?
well don't keep us hanging... (what was it?)