I suppose 'chance' is a little hand-wavy, but isn't a p-value just the probability of your data given that your hypothesis is false? Isn't that literally and precisely the probability that they occurred by chance?
I suppose 'chance' is a little hand-wavy, but isn't a p-value just the probability of your data given that your hypothesis is false? Isn't that literally and precisely the probability that they occurred by chance?
No, it's the probability of a particular observation, given that we assume the result is due to chance. This sounds similar, but the difference is that it doesn't say anything about the probability of your result outside the context of the study's hypotheses.
It's the probability of you seeing your results due to chance if there is no effect.
This differs from the probability of the results being due to chance because it does not take into account the probability of your hypothesis being true or false.
If you observe something that would disprove e.g. General Relativity with a P of 0.001 it is much more likely to be due to chance than if you observe something that is consistent with known science with a P of 0.001, as the weight of all the evidence for General Relativity is very strong.
> Is the p-value really not the probability of your results being due to chance?
No, it is not. It is the probability that the statistical attributes of the data would be equal or more extreme than observed assuming the Null Hypothesis is true.
In your definition, there is no assumption of what model the data should conform to...so what does "by random" mean in that context? Also, by random doesn't mean 'not predictable' or useless. If I roll a 100 sided die an infinite number of times, I'd expect the number of observances of '5' to approach .01 of the total distribution. So, the probability of rolling a 5 by random chance is 1 in 100. However, I would not reject my Null Hypothesis since my model predicts exactly this random behavior.
Now, if I rolled a die 1000 times and rolled a 5 every time the mean of that distribution (5) would be very, very far from the expected mean of my model if the Null is assumed to be true. And I may be tempted (very) to reject the Null that I am rolling a 100 sided fair die.
I will now sit back and wait for my definition and analogy to be torn to shreds. :)
In conditional probability notation, it's the difference between P(result | it's just chance) and P(it's just chance | result).
You can't actually say unless you either (1) roll the die more times, or (2) assume something about the probability that I gave you an all-7's die to begin with.
Doing (2) is useless, because that exactly the question we are trying to answer.
For example, suppose I perform this experiment all the time and I know that I give an all-7's die only 1% of the time. With this new information, you could actually calculate the probability of an all-7's die given a 7 roll. Of all the possible outcomes, you could add up the ones where I gave you an all-7s die and the ones where I gave you a normal die but you just rolled 7. Then you could divide that by the total number of possible outcomes.
But this would give you a totally different number than if I give you an all-7's die 99% of the time. And the problem is that you don't have any information about what kind of die you have before you roll it. You're trying to figure out which world we live in -- one where your hypothesis is true or one where it's not.
(I am pretty sure that what I wrote above is true. But one thing I'm not as clear on is how multiple rolls of the die actually can establish confidence percentages. How many rolls does it take to actually establish confidence? Would love to hear from any stats experts about that.)
However, if you can bound the prior probability on the low end, you can make meaningful answers. Let's say you think there's at least a 1-in-a-billion chance that you have an all-7's die. After five rolls in a row of 7, there's at least a 0.3% chance that you were handed an all-7's die. After six rolls there's a 6% chance, and so on.
Usually this quantitative analysis isn't formally done, since the priors can always be debated, but rather a very small P value is demanded for very unlikely events.
It is the probability of achieving a result at least as far from the hypothesized value exclusively due to random variation that is uncorrelated with the explanatory variable(s) at hand [and subject to a number of other assumptions].
Yes, but the problem is that it also includes more than that. One of the many problems with p-values is that people assume that the p-value encodes a lot more information than it actually does.
Another problem: it doesn't really tell you any "new" information. All equality-based null hypotheses are false, and we know this due to continuity theory[0]. Really, all a hypothesis test does is tell us if the sample size is large enough to reflect this knowledge. Literally any null hypothesis can be rejected, as long as the sample size is sufficiently large.
It's worse than that, though. Yes, hypothesis testing doesn't tell us anything about the practical significance of the results. But it also doesn't actually tell us what most people thing it does: a measure of how wrong our hypothesis is. Rejection is a binary state: there's no concept of "strongly rejecting" a null hypothesis[1].
[0] We could reject continuity on the basis that the real world is actually discrete at the quantum level, but a lot of the math underpinning hypothesis testing falls apart if you don't assume continuity, so it's broken either way.
[1] Many people - including statisticians - will sometimes imply this colloquially ("the p-value was .00001, so our hypothesis was completely wrong"). It's not always "wrong", because largely you can measure this concept in other ways, but you cannot do so with a p-value.
Roll a dice 100 times, and on average 5 rolls have the pattern I want.
v.s. I have found pattern and there is a 5% change it was due to random.
Is that not the same?
* Ok, a chosen null hypothesis, often chosen to be something like "chance".