P values are not as reliable as many scientists assume
nature.com
nature.com
Now I refuse to use p-values and deliberately construct analyses that are incompatible with Fisherian statistics. And rather than giving people raw numbers, I produce a massive document of interpretation. Takes a huge amount of time, but I'm hoping it will mean my publishing track will contain significantly (ha!) fewer false results than most biologists'.
Attitudes like this bug me so much. From my experience it is a big part of what makes people that are not familiar with science not trust scientists.
It doesn't matter how rigorous scientists are with their methodology, my 'new age' mother isn't going to stop believing in fairies in the trees because scientists say there's no proof. Or when she talks about higher and lower vibrations, she is puzzled when I ask her "vibrations of what? what is it that's vibrating?". This isn't important to her story, and is merely a detail that makes life less enjoyable, so she ignores it. A less obnoxious version of ICP's "scientists are ruining life by explaining things" during their 'how do magnets work' phase.
Saying that people don't trust scientists because of p values is a bit like saying people don't trust plumbers when they use blue plumber's tape instead of gray. People don't like plumbers because they don't show up on time and charge through the nose - things that change your own story.
http://www.inference.phy.cam.ac.uk/mackay/itprnn/ps/457.466....
I have been working on a guide for practicing scientists which attacks common statistical misconceptions:
http://www.refsmmat.com/statistics/
It's on its way to publication, so you'll be able to hit your biologists upside the head with it in due time. (Although it'll be slim, so it won't hurt too much.)
---
Imagine everyone in the USA gets sudden amnesia. We want to find out who the President is, but no one can remember.
A scientist comes up with a test to determine if someone is the President.
If they are the President, there is a 100% chance the test will say they are the President and a 0% chance the test will say they are not the President.
If they are not the President, there is a 99.999% chance the test will say that are not the President, and a 0.001% chance the test will falsely say they are the President.
Giving the test to the person sitting in the big chair in the Oval Office is useful, because it's already quite likely this person is the President. If the test is positive for Presidency, it's extremely likely that person is the president.
Giving the test to the 10 people nearest the oval office is useful, because it's fairly likely the President is one of these people. A positive result will indicate strongly that that person is the President, and if no-one in that group is actually the President, there's a 99.99% chance the test will say so.
Giving the test to the 1000 people in the White House is pretty useful, because it's pretty likely the President is in the White House, and if none of these people are the president, there's still a 99% chance the test will be correct. A positive result for any one person will indicate quite strongly that that person is the President.
But giving the test to everyone in America is not very useful at all, because it's very unlikely that any particular person is the President, and we can expect the test will give a positive result for around 3200 people. For any particular person in this group, it's much more likely they're not the President than they are.
---
Is this a broadly correct, if non-rigorous, analogy? I realize most HNers will be much more familiar with this stuff than I am, I'm interested chiefly in whether or not I misled my friend.
http://www.math.hmc.edu/funfacts/ffiles/30002.6.shtml (sorry its a bit informal, it was first google hit, but I have seen that in real textbooks for sure)
Although its a bit confusing associating the prior with geometric proximity to the oval chair.
If you test one person and the the test is positive, then that person is the president (p=0.00001).
If you test a thousand people, and the test is positive for one of them, then that person is the president (p=0.01).
So you don't really need Bayesian logic to reason that you should test fewer people if you want a more significant result. (Note I'm not saying you don't need Bayes' Theorem, which everyone uses.)
Edit: I think most people on HN get their knowledge of frequentist and bayesian statistics from XKCD #1132. That's sad.
That's misleading by the use of the word "significant" which apparently means something in frequentism than it does in normal speech. I certainly wouldn't use "significant" in that way as a non-frequentist, I would instead rephrase what you said as:
> So you don't really need Bayesian logic to reason that you should test fewer people if you want more confirmation bias in your result.
And that's a statement I can definitely get behind!
Replication and plausibility in the context of other studies are taken into account, however these additional parameters are difficult to reduce to numbers.
A recent blog post (link below) describes the actual practice of (good) biomedical research fairly well.
http://www.tcpinnovations.com/drugbaron/the-all-new-good-old...
I am frequently a frequentist, but this particular test doesn't make sense to me. Without any prior belief, the person you tested is still no more likely to be the president than any of the other 3199 who would test positive.
The ~99.97% chance that the test result is a false positive for a subject chosen at random is a consequence of the prior probability of that subject being the President in the first place being over 3000 times lower than the probability of a positive test result. In other words, though the sensitivity of the test is perfect, the specificity of the test is insufficient to isolate a condition as rare as the Presidency. It has nothing to do with testing one person being inherently more decisive than testing more than one.
Only narrowing the pool of candidates in a way which will increase the prior probability that the subject is the President will improve the situation, not random decimation to a single test subject. The GP gave several correct examples such as limiting the test to those who were much more likely to be the President in the first place (eg. a person randomly found sitting in the President's chair at the Oval Office might have a prior probability of being the President of greater than 0.01, millions of times greater than that of a randomly selected person in the US).
You've got it backwards. The p-value is the chance that we see the given data given that X is not president.
In your explanation, where you said "given a sample" you should say "given a sample size", to make it clear that the type-1 error probability is not conditional on the sample data. If you condition on the observed data, then the probability of passing a statistical test is going to be 1 or 0 depending on whether or not the data passes the test. It is the test itself, not the error probability, that depends on the sample data. Which is also something you should have specified.
That is, if I test 1000 people at a high school football game in Peoria, Illinois, and one of them comes out positive (p=0.01), there is a very different likelihood that I have found the true president than if I test 1000 people at a official dinner in the White House and find a match (also p=0.01). In fact, I think I'd be willing to wager more on the chance that the individual tested in the White House is actually the president than the random football fan in Peoria, even if only a single person in Peoria was tested (p=0.00001).
Of course, this is only an issue if p-value is being used to express relative confidence in a conclusion, which it shouldn't (?) be. Still, how does a frequentist account for a choice of venue like this?
A frequentist usually builds a more descriptive model by adding more variables (accounting for obvious factors that contribute to differences in observed frequencies) and increasing sample size to increase statistical power (reduce false negatives). This is not applicable in this case because there is only one president and the tests have low power. The test is bogus because you can't 'sample' and determine to probability of X==president effectively.
In this case, the experiment is really "who is president?" and not "is person X president?" The "who is president?" question does not really have a null hypothesis, so it makes no sense to talk about p-values. Instead, we can talk about two different results: identified the president correctly, and identified the president incorrectly.
We then have to design the experiment in such a way that probabilities can be calculated. But this is not possible if you say the repeated experiment is testing a bunch of people to see if one is president--because the probability depends on who actually is president, and that's not known.
So you consider the experiment to be "the president goes about their business and ends up in some random place, we get amnesia, and then run the test." We can simulate this experiment because we can come up with some probability distribution for where the president is. We can't come up with a probability for who the president is, because that's unknown and either 0% or 100%--only Bayesians would let you do that.
Then you can choose the order of people to test to maximize the probability that the president will be identified correctly.
And then you say things like,
* The test had a 92% chance of succeeding.
* We only had to test 14 people, and in such cases, the test had a 98% chance of succeeding.
But even if they published them with the verbiage you mentioned, that would be a better world because reporters would be less likely to turn it into a story with a headline like "Scientists say ...". The prevalence of articles like this erode trust in science, and fill the public's head with misconceptions.
If you were equally mocking of both Bayesian and Frequentist, the Bayesian would arguably come up with basically the same conclusion. In this scenario where the Frequentist ignores previous data based on human history and our understanding of the lifecycle of the sun based on physics, the Bayesian should do the same if we are fair. Then his prior distribution would likely be 50% belief that the sun would explode (Why favor either outcome? Thus the prior distribution is split down the middle between both outcomes.). With the new information given from the detector, his posterior would be 35/36 probability that the sun has exploded.
Maybe the Bayesian approach provides a more explicit way of incorporating previous information, and without that, it's cause for some misuse with a Frequentist approach, but that doesn't mean Frequentists need be fundamentally ignorant in their approach.
The comic's good for a laugh. Just don't take it too seriously as a criticism.
So just to get a handle on this stuff, the two problems you have if you do a test and only look at a low p are:
1) Unusual things do occur. If a million people do the same test, it's obvious they'll come up with some wrong values. It's less obvious that a similar number of wrong values will come if a million people do a million different tests with a similar small chance of bogus results.
2) The pattern of results may indeed be unusual but not necessarily in the fashion you think it is. There may be a non-random pattern possessed by the data but may not because of your particular hypothesis but a "this is not random" result may seem to say your hypothesis does explain the data.
Does that characterize the problem?
Say you are looking for genes that might influence the rate of occurrence a particular disease. There might be genes that influence this rate, or there might not, it could be entirely environmental, or it could be entirely genetic, or something in between. In any case, you go genome-wide studies, and find that certain gene variants occur more often in your diseased population than in your control population. You apply frequentist statistics, using some corrections for multiple hypothesis testing, and get some kind of "significant" result. This gets published in Nature (you lucky thing!).
Are your conclusions correct? Do the genes you identified really modify the course of the disease you studied? Bayesian statistics won't give you the answer.
The only way to get the answer is to do experimental science, i.e. deliberately modify the gene(s) in question and show that your modifications change the occurrence or course of the disease.
Unfortunately, that is not always feasible, for either technical or ethical reasons, so we have to fall back on the poor cousin of experimental science that is population statistics.
Reminds me of this quote: (These values are known as priors, which is ironic because Bayesians inevitably pull them out of their posteriors.)
Last week I did some Bayesian (!!) work on finding a pdf for the direction of a certain quantity. My priors were... [0,2*pi).
If you are hypothesizing that genetics influences the incidence of a disease, without any prior knowledge as to whether there is any influence of genetics on a disease, how can you have a prior probability that there is a genetic influence on the disease in question?
- There are more advanced tools in Bayesian analysis such as Jeffreys prior (known as uninformative priors, look it up).
- As was mentioned in other responses the same "big problems" exist in every other statistical and mathematical modeling approach, namely that you have to make assumptions and your results are going to be crap if your assumptions are crap.
- Generally, Bayesian stats got a late start due to high computational resource costs, not some theoretical limitations. The issue with priors that gets repeated by philosophers and some statisticians does not stop the huge, monumental progress Bayesian statistics has had in a ton of applied fields, from computer science / machine learning all the way to economics and political science.
Edit: language
P(gene|healthy) = 0.010 and P(no_gene|healthy) = 0.99
P(gene|sick) = 0.011 and P(no_gene|sick) = 0.989
Then you do large population studies (testing for sick vs healthy is cheaper than testing for genes). You get P(healthy) = 0.90 and P(sick) = 0.10.
From the small sample you get the average of the gene's expression frequency: P(gene) = 0.0101 and P(no_gene) = 0.9899.
You are interested in what is the probability of a person being sick when they have the gene P(sick|gene) =
P(sick,gene) / P(gene) =
P(gene|sick) * P(sick) / P(gene) =
0.011 * 0.10 / 0.0101 =
0.109
And conversely, odds of being sick without the gene: P(sick|no_gene) =
P(no_gene|sick) * P(sick) / P(no_gene) =
0.989 * 0.10 / 0.9899 =
0.0979
So, based on this data, having the gene increases your probability of having the disease some 11% (0.109/0.0979) or 1.1 percentage points (0,109-0,0979).Or you might want to calculate some other figures from these. If you just want the ratio, then P(sick) does not have a bearing on it. You could perform some sensitivity analysis etc...
(I am not a statistician by far)
P(gene|healthy) = 0.010 and P(no_gene|healthy) = 0.99
P(gene|sick) = 0.011 and P(no_gene|sick) = 0.989
My point is that in many cases there is no justification for placing any number whatsoever on a prior hypothesis. You can't simply say that the prior probability of a particular gene being involved in a disease is 1%, or 0.00001% or 10%, or whatever.
Edit: I'm not saying that Bayesian statistics is without uses, it is very useful in epidemiology, for example. However it is not appropriate for determining molecular and genetic mechanisms.
P(gene|healthy), P(gene) and P(healthy) are something we can often measure.
P(healthy|gene) is something we can calculate with the Bayes formula from the above values.
More generic version:
We want to know P(model|data). That's what we always want. What is the probability of some model, based on this data that we measured.
But we only have P(model), P(data) and P(data|model). So we use the Bayes formula to get the answer to the interesting conditional probability. It needs all three inputs. We must estimate if we don't have some.
Just presenting P(data|model) and not constraining P(model) or P(data) in any way means that we can't say anything about P(model|data).
It's like if we start with x + a + b = 5.
If we don't know or estimate a and b, it's impossible to say anything about x. Bayes' formula is like saying then, solve it like this: x = 5 - a - b.
So, if you say, there is not justification placing any constraint on a or b - then it's just saying we clearly do not have enough data to say anything about x either. There is no way out of it.
And in those cases, frequentist statistics is no better off, because you have no sense of what the sample space is.
However, there are cases where there is no well-defined sample space, but you can still assign reasonable priors; so Bayesian statistics covers a range of cases that is a superset of the range of cases that frequentist statistics covers. E. T. Jaynes goes into this in some detail in his book Probability Theory: The Logic of Science.
It's a poor analogy, because it's not clear to people why such a test is "natural." It's not clear how your specific test could be broken in the peculiar way that it would have 99.999% chance to confirm that someone who isn't the president is indeed not the president.
And people would get caught up with what you mean by "If they are" and "If they are not," since it's not clear how you would know the error of your test without a real president around to identify.
False positives or false negatives are not at all intuitive to people who have never done experimental design. Most people would get stuck at percentages anyhow.
They would muddle my otherwise irreproachable statistics.
Of course, TONS of people forget this and publish a p-value as if those two variables are the only ones under consideration. Which is just sad.
P values have always had critics. In their almost nine decades of existence, they have been likened to mosquitoes (annoying and impossible to swat away), the emperor's new clothes (fraught with obvious problems that everyone ignores) and the tool of a “sterile intellectual rake” who ravishes science but leaves it with no progeny3. One researcher suggested rechristening the methodology “statistical hypothesis inference testing”3, presumably for the acronym it would yield.
"One researcher suggested rechristening the methodology “statistical hypothesis inference testing”, presumably for the acronym it would yield."
I'm continually surprised at how many people either don't know, or don't internalize, that. Look at how often "risk factors"--which are a descriptive concept--are converted to advice--which is predictive.
Doing so in the absence of a causal hypothesis is a basic violation of "correlation does not equal causation."
If you want to construct a scientific theory you must be able to articulate some predictive tests, and that means you must hypothesize a causal mechanism.
Does the law of the large numbers somewhat unify descriptiveness and predictiveness?
Correlation suggests that if you're looking for causation, it might be somewhere over here. It doesn't insist that the two are the same, but if you're looking for clues, it's a hell of a dowsing rod.
1. The field, in general, demands far more mathematical rigor when dealing with statistics.
2. The demand for mathematical rigor is justified because most data sets we deal with are many orders of magnitude larger than what psychologists and others encounter. So predictions based on limit theorems, etc. are often testable.
I've gotten a fair amount of "I just need a p-value" requests, and some assumptions that everything can be hit with a t-test or an ANOVA and it'll all work out fine.
If you read calculus with about the same fluency as the comic books then "Data Analysis: A Bayesian Tutorial" is awesome http://www.amazon.com/Data-Analysis-A-Bayesian-Tutorial/dp/0...
And if you would like a little more exposition (but still a mathematically sophisticated treatment) "Doing Bayesian Data Analysis: A Tutorial with R and BUGS" is fantastic http://www.amazon.com/Doing-Bayesian-Data-Analysis-Tutorial/...
The latter will also give you more details of how to approach classical, frequentest tests and summary statistics with their Bayesian equivalent.
Honestly I would say get both books as they're cheap and provide different insights. You only need to read a few chapters of each to see how you approach basic experiments from a Bayesian perspective.
Simonsohn, has a whole website about "p-hacking" and how to detect it.
He and his colleagues are concerned about making scientific papers more reliable. You can use the p-curve software on that site for your own investigations into p values found in published research.
Many of the interesting issues brought up by the comments on the article kindly submitted here become much more clear after reading Simonsohn's various articles
http://opim.wharton.upenn.edu/~uws/
about p values and what they mean, and other aspects of interpreting published scientific research. He also has a paper
http://papers.ssrn.com/sol3/papers.cfm?abstract_id=2259879
on evaluating replication results with more specific tips on that issue.
"Abstract: "When does a replication attempt fail? The most common standard is: when it obtains p>.05. I begin here by evaluating this standard in the context of three published replication attempts, involving investigations of the embodiment of morality, the endowment effect, and weather effects on life satisfaction, concluding the standard has unacceptable problems. I then describe similarly unacceptable problems associated with standards that rely on effect-size comparisons between original and replication results. Finally, I propose a new standard: Replication attempts fail when their results indicate that the effect, if it exists at all, is too small to have been detected by the original study. This new standard (1) circumvents the problems associated with existing standards, (2) arrives at intuitively compelling interpretations of existing replication results, and (3) suggests a simple sample size requirement for replication attempts: 2.5 times the original sample."
0.05 p-value means that there is 5% probability that (for a t-test as example) a difference in averages of two sequences (the statistic) is by chance and not because of difference in means of their underlying normal distribution.
I assume that the "toss-up" means that there is no difference in the means in reality (so the null hypothesis is true). Am I understanding it correctly? Shouldn't in this case the probability of getting p-value < 0.05 be in fact less than 5% and not 29%?
How did they get the 29%?
I'm not 100% sure how it works, and I've never seen it applied in biology, but this is medicine.
Do you see the difference. The p-value doesn't tell you anything about the population, it only gives you information about your sample -- a confusion that this article was pinpointing.
So what the 29% is telling you is that, the chances of finding a statistically significant result (p = 0.05) given a population factor of µ = 0 are up to 29%. Whereas given a population factor of µ = 0 the chances of getting a significant result (p = 0.05) is 5%.
EDIT: erased a wrong and confusing claim.
EDIT: also, what is population factor µ? The difference in means?
First, the reported p-value might be wrong. E.g. basing it on assumptions of normality when the data is non-normal. However modern non-parametric approaches like the bootstrap can avoid this issue.
Second, testing multiple hypotheses. If you test 10 hypotheses then you cannot reject the null (that all 10 null hypotheses hold) simply because one single hypothesis is rejected in isolation. But this is well known, and failing to account for it is an issue with the researcher, not with frequentist statistics. I actually think that the main practical difference between Bayesian and Frequentist statistics is whether accounting for the issue of multiple hypotheses is done formally or informally.
You're absolutely correct about using non-parametric tests, and more scientists should be using them. The normality assumption is flat out laughable when using real-world data most of the time.
You're also correct about multiple hypothesis testing. Accounting for familywise error (e.g., Holms adjustments) can help to keep your p-value reporting honest.
That doesn't negate the underlying problem, though. A p-value is simply an indication, nothing more. The p-value never promised to be more than that. The issue isn't in the p-value's construction, the issue lies in its misuse and how easily it can be abused in statistical reporting (see: p-hacking).
The p-value as a test statistic is perfectly honest in my opinion. But like many other statistical methods, it comes with its own set of baggage that I feel gets conveniently glossed over more often than it should.
What is the HN recommandation?
More information is better.
Effect sizes and confidence intervals are much more useful than a p value or two.
Does it work? Sometimes.
The problem is that people tend to believe something they use a lot.
Even the 0.05 threshold is sort of made up.
Correlation does not mean causation.
the normal distribution is also quite well justified by the central limit theorem.
i do however agree that a p-value of 0.05 is not worth very much.
That's what is supposed to happen, though, right? You publish your findings. Others try to reproduce. They publish THEIR findings, etc. etc. If most published findings are false, it sounds like the process is working as designed.
Most scientists have better things to do than replicate previous findings, unless that previous finding directly bears on their own work.
consider the wikipedia example with heads vs. tails.
http://en.wikipedia.org/wiki/P-value#Examples
the idea that 5 coin tosses can produce a p-value < 0.05 that 'demonstrates' that the coin is biased towards heads is intuitively 'obviously wrong'. even if we take it to 10 coin tosses (the p-value you get is 0.001 - which looks really strong if we accept that 0.01 is acceptable) it clashes with my own ideals for what statistical significance should mean. this is in a loose way a proof by contradiction that p-values of 0.05 or 0.01 do not have utility (at least for these kinds of small n).
aside from that consider running the experiment 5 times or 20 times. how many false positives do you expect? what is the expected number of false positives? is that significant?
it also bothers me how connected to the problem formulation that the value itself is. if we analyse the same situation with an identical test but a different formulation of the problem that the values differ?
why is five heads in a row less significant as a result when the test is whether a coin is biased at all rather than a test that it is biased towards heads only? sure i understand the probability involved there that we have all these potential coins biased towards tails that mean nothing in the first case - but there is something very deeply wrong with that.
shouldn't this be the other way around? if 5 consecutive heads is good evidence that a coin is biased towards heads, isn't it equally good evidence that it is biased at all? classical logic says that it is because being biased towards heads is a subset of being biased in either direction. the truth is that it really is equally good evidence - i challenge someone to explain why it is not! ( actually i kinda want to be wrong about that because i might learn something new then :) )
probability is counter-intuitive and useless for the kinds of small n usually used in experiments - the intuition about it recovers when we deal with sensible n - numbers like 1000 or 10000 - but these are still small n really if you need to scale up, or be confident that your result is correct. even at 100 samples its obvious that our idealisation of percentage and what happens in reality do not marry up neatly...
to make a very crude software analogy what about those 1 in 10,000 bugs? they are still a very real problem if you have millions of customers...
or - IMO even 10,000 is a very exceedingly small n to try and draw robust conclusions from.
You should really get used to the idea that stating a different problem will give you a different answer. You need to be very careful when asking a question, or your answer might not mean what you think it means.
> You should really get used to the idea that stating a different problem will give you a different answer
this is entirely normal and expected since ever i can remember... what i'm saying is that you can analyse the same data in two different ways and reach differing conclusions because of the nature of the p-value (vs. the nature of, well... nature)
what i was mainly trying to get across is that a coin being biased towards heads /logically infers that/ it is biased. so the idea that 5 heads in a row is less evidence of a coin being biased than it is that it is biased towards heads is not only counter-intuitive but in disagreement with a much stronger and more intuitive form of reasoning.
the fact that the p-values are different in these cases leads me to expect that p-values on their own are not a good indicator of strength of evidence without a lot more context - and /really/ understanding what that context is and means - in which case why use the value at all? nobody else is likely to interpret it correctly unless you lay it out that way which then negates the supposed utility of the p-value...
and yes, i intuitively consider 5 heads in a row to be unspectacular for a fair coin certainly not a 19/20 chance that it is biased (maybe i am very, very wrong though).
The evidence is different because there are different outcomes. For the two-tailed test the possible outcomes are: biased toward heads or tails, or not biased at all. For the one-tailed experiment, the outcomes are: biased toward heads, or not. Getting 5 tails in a row would be evidence in favor of the coin being biased, but not being biased toward heads.
Think of it this way: the two-tailed test is running 2 experiments at the same time (one for heads and one for tails) with the option of picking the one that gives you better results. So obviously the standard for significance has to be higher, because you're cherry-picking results. https://xkcd.com/882/