It seems like something that should be taught alongside math from elementary school, not something you can maybe get an elective in in high school or college.
I actually found statistics pretty fun and eventually we learned to use R for it. Though I ended up missing 1 lecture and it almost screwed me over for the rest of the class. I found a Kahn academy video though that covered it and was able to catch up. There really is a ton of information to learn. It would have been nice to have at least touched on it a little bit in highschool beyond learning about means and the tiniest bit about standard deviations.
I know a lot of people have trouble thinking clearly and correctly about probability and statistical inference.
But do we know if, practically speaking, that can be addressed by a modified educational curriculum?
Or are these concepts that would take an extraordinary amount of effort for many persons to understand well?
[1] source: I never saw it as a student, and I've consulted with math teachers (well, administrators, about math) in multiple states, and I've never had it come up there either.
Everybody wants to compute p(H|x), the probility of a scientific hypothesis given the data. People want to do this so badly that they can't help interpreting the p-value that way.
You can actually compute p(H|x) if you use Bayesian stats.
Typically people doing research have prior information about what they are researching. E.g. previous studies have found effect sizes to be in some interval.
You can fallback on an uninformative prior or other tricks if you really want to model the idea that you know literally nothing about what you're researching. But that should be very rare.
Our 0.05 p-value limit is effectively just answering that question with a fixed 5%, no matter how ridiculous the proposition is.
You can't "compute" it for any useful meaning of the word "compute". You can estimate it intuitively, or you can try to look at how many pre-registered unpublished studies or null results have been published. Otherwise there's no way to even consider getting a grasp of what P(H) would be.
That's called picking a prior. Usually, pick the prior that maximizes the entropy on the given interval (zero prior knowledge).
For those reading and confused. Computing: p(H|x) = p(x|H)p(H)/p(x)
requires p(H) which the parent is suggesting is impossible to grasp.
If we want to know the probability that a coin is biased, we can assign probabilities to each hypothesis. For some people, p(bias=1/2)=1 (Dirac distribution), others might argue that p(bias=x)=1 for x in [0,1] (uniform distribution). Others might argue its some Beta function centered around 1/2. I believe what the parent is suggesting is that choosing which original belief we have in the system is a matter of philosophy, not computation.
Essentially, we want to get p(H0|x), but we need Bayes Law to get this from p(x|H0). But we need some notion of what priors to use. This is of course impossible to actually get, but if we published null studies then it would allow us to estimate it with better confidence. Alternatively we can ballpark it by saying how unexpected the result is, and how long the effect being explained has been studied.
The "common-sense" version of this is intuitive, that "extraordinary claims require extraordinary evidence". The more unexpected a result is, the lower the p-value has to be to be convincing.
The epidemiological version of this shows up in things like the Bradford Hill criteria [2], which include significance of association but also attempt to bring in plausibility.
[1] https://statmodeling.stat.columbia.edu/2015/03/02/what-hypot... [2] https://statmodeling.stat.columbia.edu/2015/07/03/why-should...
The problem is that we are imbuing the words "statistical significance" with a whole bunch of math. This would be fine, except for the inconvenient fact that people also want to use the word "significance" as it is defined in English.
It is not only possible, but likely that people will be producing results that are insignificant but statistically significant. I mean seriously, if I tell my boss that the results are statistically significant, how do I expect him to understand that the results might reasonably be insignificant? How do I expect anyone non-technical to take that sentence seriously? You lose all credibility pretty quickly when people start saying that the result is significant except it doesn't matter.
This is even more stupid than calling complex numbers "imaginary" and "complex". Those names are just arbitrary. Statistical significance is overloading words that are likely to come up with meanings other than what they are expected to mean.
Anyway, this is the problem that things like this article are going to. People are treating significance values as though they are significant. They aren't significant, they are statistically significant. It means something different to significance.
For example; if I run an experiment 20 times and get a p value consistent with a 5% chance in one experiment, it is statistically significant but the result is obviously not significant in the English-language meaning of the word.
In fact, failures to reproduce "landmark" experiments have low significance (statistics word) because they find a result that is highly probable if the null hypothesis is true. But failures to reproduce have high significance (colloquial English word) because they indicate that widely-held beliefs are wrong. So in some of the most important cases, the statistics and colloquial English meanings of the words are exactly the opposite.
I don't know if there's a good solution to this, though--it's hard to change organically-emergent trends in language even if we did come up with a better term.
If you look at research on carcinogens, it is frequent to see something like "it is statistically significant that eating chemical X increases your risk of cancer" but when you read on, the magnitude of the effect is "increases risk by .0023%, for a cancer that already has a .6% incidence in the population".
My point is that people who are not statisticians (or people who use statistics), tend to be concerned with the magnitude of an effect, and frequently misconstrue the statistical significance for that magnitude.
But these concepts aren't clear to the people using the data. I've seen people use statistical tests done in batch without correcting for multiple testing (one of my regrettable tasks), and flat out ignore any differences that weren't statistically significant. They thought it meant "no difference."
With population-level data (e.g., Census counts or hospital records), I sometimes wonder if statistical significance has any value. It's a tool to make sure we don't accidentally claim a difference where there isn't one. But I can guarantee you that nothing worth measuring is exactly the same for men, women, whites, blacks, young, old, whatever. It'd be amazing if the difference was 0. So statistical significance ends up being an obfuscated way to state the sample size.
The problem is that people want an answer to the question, "Here is a pile of data, what should I believe?" But mathematically, the proper role of data is to modify existing beliefs, and not to dictate beliefs.
Statisticians can spend forever explaining this. But instead have gone with the cop-out of asking a question that is confusingly similar to the one that people want to ask. It is popular exactly because it is so easily misunderstood. You can give any number of lectures on what it actually means - I guarantee that it will be misunderstood.
Worse yet, p-values are sensitive to particulars of experimental design that logically should never matter to inferences. My aunt and uncle provide a classic example. They wanted a son and a daughter. They had 6 sons then a daughter. Are they biased towards one gender? The p-value for this question is .03125 (we'd see this much evidence against with all boys, all girls, last girl, last boy, so 4/2^7). But if I instead said that they planned to have 7 children, and had 6 sons then a daughter, the p-value changes to 0.125 because the one child out could have been anywhere in succession. But in Bayes' formula, their plans if something else had happened cannot ever affect how we adjust our inferences.
(This example is not made up. When Lorna found that they had a daughter, she told Bill to get a vasectomy from the delivery table. By another crazy coincidence, their 7 children were also born on all the days of the week.)
Its not just about p-values or the likelyhood of making a particular type of error by agreeing with a thesis.
There is a fundamental misunderstanding that most people, and even many scientists make with regards to the philosophical interpretation of 'truth' using the scientific method.
By definition, scientific 'truth' is always fallible, and the vast majority of people have a very hard time dealing with this concept. A better explanation or additional data, or better data is always possible and needs to be possible for science to not be dogma. This is fundamentally different than how people interpret truth and understand what the word means. All scientific truths are replaceable.
That's not how statistics works. Frequentist p-values and Bayesian inferences are both entirely dependent on the mathematical model in question. "Their plans if something else had happened" are a key part of the model. What you're describing at 2 entirely different experiments:
1) A couple has children until they have 1 of each gender. they have 7 children. How likely is a result this extreme (or greater)?
2) A couple has 7 children. They have a gender ratio of 6:1. How likely is a result this extreme (or greater)?
The probability of event A given that we observed B is the probability of A and B happening divided by the probability that B happened. Mighta, coulda, shoulda but didn't doesn't enter into it and can't affect the result. The fact that frequentist statistics does care and shouldn't is one of the major criticisms that Bayesians offer.
If you think you understand statistics and don't understand this fact, then you do not understand statistics as well as you think you do. But you can be pardoned. Most statistics classes are too busy cramming statistical tests into student's heads to bother them with bothersome facts about where the cracks in the foundations are.
So true. That said, here's a great interactive explaining the p-value: https://www.jwilber.me/permutationtest/
The big problem is that since we've already long since entered the world of big data you can now find spurious correlations, similarly strong, but constrained to variables that are plausibly connected. For instance what if this correlation was between spending on science/space/technology and an increase in the number of people pursuing STEM fields? People would immediately just accept it without question, even though it should in theory be held to the same scientific rigor and critique as one that challenges our biases. Science should not be a glorified exercise in confirmation bias. Spurious correlations don't mean you have a biased sample or that such correlations don't exist - it simply means that correlations have no direct relationship.
538 has a great little interactive p-hacking game you can play, with real data. [2] Ultimately any single datum that tries to legitimize a piece of research is going to be gamed. Science should be judged as it was in its 'glory days' - by the logic, predictability, and falsifiability of the ideas presented. Anything that steps outside these bounds should be treated with the most extreme of prejudice. It may be legitimate but the great burden of proof is on the presenter, and that doesn't come down to p-hacking up a 0.00001 correlation and calling it causal.
In my work as a data scientist, 99% of the time this is not even relevant. We have huge sample sizes and most of the p-values we see are <.0001, meaning there is SO much data that even a very very small effect size can be found "significant."
The question then: does a mean of 20.02 being significantly different from a mean of 20.03 over 10 million patients, is that actually important to us?
[1] http://computationalimagination.com/article_practical_signif...
You might see p=0.00001 but if the size of the effect irrelevant, then the statistical significance of the relationship is still not something I care about. Great, orange juice shrinks tumors by 0.000001%. I'm still going with chemo, thanks.
People that work with computers often are slightly better about reasoning about probabilities because they often deal with high frequency events. A 1% error rate across 10 million events a day leads to 100,000 errors. If those events are transactions, payments, file uploads etc, a 1% chance starts to look very common. But even computer people are not immune to incorrectly reasoning about probabilities. I suspect there is some fundamental limitation at the cognitive level.
In my view, understanding p-values won't help. Feeding data into a formula that produces a p-value doesn't account for things that can go wrong with experiments: Uncontrolled experimental conditions, biased sampling, hidden correlations, and simply non-ergodic systems. It is possible that some systems can't be controlled to the point where it's reasonable to start looking for real effects.
I suspect that if the results were any good, the precise interpretation of the p-value wouldn't matter. Physics has lived without an agreed-upon interpretation of quantum theory for a century.
My decision to select, implement, go, or no go cannot itself be a probability distribution.
If we required scientists to be good statisticians there’d be far fewer scientists.
P hacking isn't always done on purpose. It's misunderstanding. That's probably why it still passes peer review. There's also tons of incentives that encourage this behavior. These are the problems. An arbitrary mark to meet makes people weak at statistics not understand their data as well. I'd rather more accurate data than more scientists. More workers doing poor work isn't useful.
Also, you can Bayes hack. Bayes helps, but it doesn't address the underlying issues.