Psychology Journal Bans Significance Testing
sciencebasedmedicine.org
sciencebasedmedicine.org
I find that there is a trend of associating "bad statistics" with "Frequentists Statistics" which isn't really fair. If you found a statistician trained only in Frequentist methods and asked their opinion on experiment design in psychological research they would likely be just as appalled as any Bayesian.
I'm a big fan of Bayesian methods, but the solutions of "we'll solve the problem of misunderstanding p-values by removing them!" is still a problem of misunderstanding p-values! The misunderstanding is the issue, not the p-value.
This situation is so bad that it merits banning frequentist for this journal and I think that's reasonable. This doesn't mean that every journal in every field should, but perhaps it will be a useful temporary measure to improve quality.
Non-statisticians have been trained using bad, frequentist methods, and one way of forcing them to retrain is by forcing them to learn new statistical tools to get published.
But if you read the original editorial, up at http://www.tandfonline.com/doi/pdf/10.1080/01973533.2015.101..., you can see that they also reject confidence intervals and Bayesian reasoning with uniform priors (which is really the same thing) without providing any guidance at all on better procedures. I fear that will just lead readers to try and guess the reliability of the data themselves, or worse, interpret the sample statistics as numbers without any associated uncertainty.
So they're doing away with poor statistical procedures, but at what cost? It's like that old joke: we've found a 100% reliable cure for cancer – bombing the planet until everybody's dead.
New methods and tools are faster to use and give better results, but they require more statistical knowledge. If the scientist applying these methods or peer review can't understand the advances, it's all for nothing. Many sciences are methodologically very conservative to the extent that it holds the science back.
How do you increase the statistical knowledge of the field so much that peer review and researchers can be expected to understand and use the new methods if they can't be trusted with p-value?
The same is true in Bayesian statistics, and even simple formal reasoning with no statistics in sight. If you make wrong assumptions, you'll get the wrong result.
The only thing you can expect statistics to do is help you change your opinion about the relative merits of opposing theories. If both your opposing theories are wrong, you will still be equally wrong.
The true flaw with frequentist statistics is that it goes out of it's way to hide this fact from you. In contrast, Bayesian stats forces you to explicitly choose a prior, enumerate your assumptions, and accept that your conclusion is based on these things.
You assume a null hypothesis that (usually) represents the status quo of no influence between the theory and the data. You then collect data. The p value then describes the probability of that data aligning with / being as a result of the null hypothesis. In other words, a p <= 0.05 says that you have <5% chance that the data came from the theory stated in the null hypothesis - that is, you have a 95% confidence that you can reject the null hypothesis in favour of your new theory.
Is that correct? I may have minced terms there because my stats training is woefully inadequate, but I think I adequately conveyed the concept?
I said the theory -- by which I meant the non-null hypothesis -- is not endorsed by a low p value (indeed, this is a major point of the article). A low p-value says "hey they doesn't look like random data" not "your brilliant hypothesis is probably true". The data might not look random because of a methodological error, outright fraud, or a confound.
This is particularly important when you consider people looking at data over and over again trying to find "an effect". Theoretically, the tests are supposed to get tougher and tougher each time you examine the data (add one degree of freedom) but in practice this doesn't happen. It doesn't matter much with large data sets, but the social sciences often use datasets where n is roughly 100, and you might only have 20 subjects in a cell.
Computing the probability that the data came from the theory stated in the null hypothesis would require a (Baysian) prior.
Also, Tloewald's reply is completely and inexorably wrong. Tloewald seems to want a Bayesian answer, which frequentist statistics can't give you.
"Given the evidence, there is a >=5% probability of the null hypothesis being true"
and
"There is a >=5% probability that if the null hypothesis were true, that your data would be at least as extreme"
The only difference I see is how you avoided saying anything about the null hypothesis, but I don't see how you can avoid saying anything about it.
if the h0 were true, then the probability of the result is unlikely, how can you not conclude that h0 is unlikely? What step are you missing other than collecting a preponderance of evidence against it?
The article never enters into this distinction. It makes it clear that people misinterpret evidence against the null hypothesis as evidence of the alternative, which is a false dichotomy.
I am confused. I also have sympathy for Tloewald at this point.
The two quantities are related to each other via Bayes rule:
P(H0|D)=P(D|H0)P(H0)/P(D)
So indeed, as P(D|H0) goes down, so does P(H0|D). But if P(H0)/P(D) is sufficiently large, you can easily have P(H0|D) high while P(D|H0) is low.
I too have sympathy for everyone confused by frequentist stats - they tend to answer the exact opposite question that one really wants answered. In contrast, Bayesian stats tend to answer the question that most people ask.
What does P(D) mean?
I read that is, the probability of Data being true.
edit: to clear up my meaning.
I mean, it makes sense to me to ask "What is the probability of getting this data, given that the null hypothesis is true"
and "what is the probability of the null hypothesis being true, given this data"
but I don't know how "this data" evaluates on its own. I can't picture that
does it mean, how authoritative is the data? Maybe that's it.
edit: OK never mind I kinda worked it out on my own.
P(D) = P(D|H0)P(H0) + P(D|H1)P(H1)Often the hypothesis testing framework is stated something like:
H0: µ = 0 (null hypothesis)
Ha: µ ≠ 0 (alternative hypothesis)
When you reject H0, it means that you can be somewhat confident that there was was some kind of distortion in your data that moved the mean away from (in this case) 0.You can create a number of theories that purport to explain this mechanistically, but you'll often need particular setups like a randomized controlled trial that can eliminate alternative explanations. When you've eliminated all the competing reasonable hypotheses, there can only be one. If you can use that hypothesis to make non-obvious predictions, that's further proof that it's right as well.
Hypothesis testing is there to tell you when to take an effect seriously, it doesn't tell you whether your explanation is right outside of very carefully constructed circumstances (i.e. where if you see a particular effect, only one theory can explain it).
For practical purposes NHST is a function that returns either H0 or Ha.
Say you have a confidence interval of 95% or higher (p value <= 5%), then the right thing to say is:
95% of the variance within the data is explained by the model (you've come up with).
I think it is ridiculous that the main focus always seems to be arcane properties of the statistical algorithm and not the answer that it delivers.
Lastly, if you're in a situation where you're performing a large number of tests, always correct for multiple testing (i.e. false discovery rate or some similar method). Also, you can take advantage of your large data set to construct a negative control (e.g. by shuffling samples, the exact method will vary) and verify that your chosen statistical test gives a flat distribution of p-values, which is the expected result when all of the null hypotheses are true. If you have an excess of small p-values in the null data set, this indicates your test is producing false positives and is not reliable (presumably because one or more assumptions have been violated).
"the problem is that, for example, a 95% confidence interval does not indicate that the parameter of interest has a 95% probability of being within the interval. Rather, it means merely that if an infinite number of samples were taken and confidence intervals computed, 95% of the confidence intervals would capture the population parameter"
And I say, "What? If out of every 100 random samples, in 95 of them the parameter is in the interval, then surely the probability of the parameter being in the interval is 95% by definition?"
Where was the catch? I remember there was one (which is enough for practical purposes because it means I won't say in a paper that the probability of the parameter being in the interval is 5%) but I feel dumb for not being able to see at first glance something that is supposed to be basic statistics...
If I have a variable that is always positive, then I could have a weird procedure to generate confidence intervals that gives me the interval [-inf,0] 5% of the time and the interval [0,inf] 95% of the time.
This would meet the definition of confidence intervals perfectly, and yet when I get [-inf,0] the real probability of the parameter being in the interval is 0%, and when I get [0,inf], it's 100%.
I wonder how large this discrepancy may be in practice (as this is obviously a made-up extreme case).
On the other hand credible intervals express `P(X ∈ [A, B] | A=a,B=b) = 0.95` (or more generally `P(X ∈ [a,b] | the data) = 0.95`). The latter is what is intuitively meant by "95% probability" of the true parameter being in the interval, because you do know a and b but not the parameter.
The example with random sampling of confidence intervals from {ℝ⁺, ℝ⁻} is indeed a good illustration of the difference.
As a completely fabricated example, suppose the true proportion in a coin flip experiment is 40%. If my confidence interval is [.45, .65], what's the probability that .40 is in [.45, .65]? It's 1. The probability that a fixed, but unknown parameter will lie in any confidence interval will be either 0 (it isn't in the interval) or 1 (it is). The _proportion_ of times the interval contains the true parameter is the level of confidence (95%).
To your always-positive example, that procedure is not particularly weird. There's always a balancing act with CI's about length and confidence level (otherwise, I could choose all reals as my interval and get 100% confidence level). That your [0, inf) interval has 100% coverage means that you could probably shrink that interval so that it has finite upper bound without losing more than 5% confidence. Hard to say without a specific distribution in mind or mild assumptions, but an application of either Markov's or Chebyshev's Inequality would allow you to make really loose bounds with only relatively minor assumptions.
In general, frequentists just don't like to talk about the properties of this particular sample, only about long-term frequencies – hence the name. Why? Because they object to the idea of probability as a degree of belief rather than as an objective measure, and given that attitude the statement that "there's a 95% probability the parameter is in this interval" doesn't make any sense: either it's in the interval or it isn't.
Huh? Unless you can bound the set of potential results, this isn't possible. Say I want to estimate the half-life of some material (bounded below, but not above). A uniform prior doesn't exist. How will the credible interval relate to the confidence interval?
well don't keep us hanging... (what was it?)
For example, EMB would say that if you have an RCT that shows that a lucky rabbit's foot works, then you have reasonable evidence to put that into practice. The issue is that, as this article points out, even with a statistically significant result for such research, there's no plausible mechanism by which a rabbit's foot makes you lucky. Therefore the RCT is just one part of the whole picture, and subject to a specific type of manipulation.
> If you run lots of phase 2 trials with different drug candidates where only a minority (lets say 10%) actually work, then with standard trial statistics (80% power and 5% false positive rate) you will get 4.5% false positive and 8% true positives – so less than 2 out of every 3 positive trial results were real. A much lower success rate than the 5% error rate commonly assumed.
http://www.forbes.com/sites/davidgrainger/2015/01/29/why-too...
This recent (open access) review in R. Soc. Open Sci. the Forbes article was largely inspired by is worth a read too: http://rsos.royalsocietypublishing.org/content/1/3/140216
The strength of significance testing is that it purposely doesn't try to tell you how likely something is to be true, only how likely the data you got was the result of chance assuming the treatment is no better than placebo. You're still taking the prior probability into account when trying to figure out the truth, you're just not putting a number on it.
My concern with bayesian approaches is that, like with frequentist approaches, the truth is still fundamentally unknowable, only now you're encouraged to put a number on that and pretend that it's science. While bayesian approaches totally make sense in trying to determine a patient's likelihood of having some disease when there is already data available for the prevalence in a population and the sensitivity and specificity of the tests, using bayesian logic to weight clinical trials strikes me as being highly dubious.
It would be one thing if SBM actually developed a framework to give a weight to each methodological feature of a trial, but so far I haven't seen much work to build a functioning system. Though if you're really honest about all the ways that you can have positive results without something actually being true, it seems like almost no amount of research will ever have a significant effect on the prior.
However, I would say that Bayesian approaches do have a big advantage in terms of helping with the interpretation problems that plague frequentist significance testing. Namely, as the OP article points out, Bayesian approaches reformulate the testing question in a way that is more intuitive, i.e. "what is the probability of the hypothesis given both the prior probability and the new data?". So yes, Bayesian methods surely do not fix everything, but since interpretation of statistics is such a major concern, they can be quite beneficial.
This is a danger sign - you are doing the same things the Bayesians do, just informally, less explicitly, and probably incorrectly.
The fact is that to make a good decision, eventually you need to compute a single number. This is an elementary fact of topology:
https://www.chrisstucchio.com/blog/2014/topology_of_decision...
That number will be based on some unproveable assumptions. That's a fact of Godel's incompleteness theorem, if nothing else. So given this, why is it "dubious" to make those assumptions explicit and obvious?
Your linked blog post states that if you make a good decision, then there is a process computing a single number which is equivalent to your process. This is not equivalent to what you claim. As a matter of fact, it's the same kind of confusion that exists around the p-value.
It's not the case that a process explicitly computing such a number automatically makes good decisions, which is what you seem to claim implicitly.
Also, Gödel has nothing whatsoever to to with this.
I don't claim you can't arrive at it by some perfect heuristic. I merely claim that you are better off being explicit about your assumptions and formalizing your reasoning. That just makes mistakes more obvious, makes your strong assumptions more clear, and makes it more likely that you will correctly update your beliefs rather than incorrectly discounting/overvaluing evidence.
You are right about godel, it's a separate theorem I'm referring to which says you need unproveable axioms. I misremembered, sorry, wrote that before my coffee.
However, I think it's important to notice that an explicit formula for your thought processes can be difficult (computationally expensive) to find. Our brains have evolved to use heuristics and "gut feelings" to make decisions, and the approach you propose forces you to throw all that away and use the much slower general purpose processing part of your brain to emulate those processes. So there's a tradeoff there.
Given that, would you support using a random number generator as part of the drug approval process to remind people of the importance of the unknown and unknowable?
Any procedure you use will have assumptions. You can't escape this. The only question is whether we show or hide them. Can you give an argument in favor of hidden assumptions and non-explicit procedures?
So as counterintuitive as it sounds, I think there are actually a couple of good arguments that can be made here:
1) With significance testing, the burden of supplying the assumptions and determining meaning is largely on the reader. With bayesian, it's transferred to the author. While it might make sense to use Bayesian for things like the Cochrane report, it's not obvious to me that each person who designs a research study and collects/analyzes data should also be in the business of trying to say whether some phenomena is real when looking at all other studies.
Essentially each study now becomes a metastudy, with all of the practical and epistemological problems that entails. The fact that it's difficult to figure out what that even means should be a red flag. (And yes, I realize this is the Chewbacca defense.)
2) So TokenAdult actually turned me onto this book Measurement In Psychology, which is all about the epistemological problems with assuming that anything you can assign a number to is a measurement. That is, having the property of being meaningful when interpreted on a ratio scale. The exact argument is kind of esoteric, but the basic takeaway is that it's very easy to trick yourself into thinking that just because you can assign a number to something that it's a measurement, to the point where assigning numbers to things in the first place tends to lead to worse decision making than if you had just used a green/yellow/red system or whatever.
As for each study becoming a meta-study, that's silly. This is indeed the chewbacca defense. Rather, each empirical study provides Bayes factors which the reader can then use to update their posteriors.
Regarding (2), obviously not every number is a measurement. In Bayesian stats, numbers representing probabilities are quite explicitly opinions. They are meaningful on a ratio scale, and are even asymptotically known to be correct. But they aren't measurements.
(They are correct if your priors are absolutely continuous w.r.t. reality. If you hold a religious belief so strong that evidence can't change it ("100% certainty"), that's not an absolutely continuous prior.)
"EMB would say that if you have an RCT that shows that a lucky rabbit's foot works, then you have reasonable evidence to put that into practice."
You don't yet know if this is a 1 in 20 result or a 19 in 20 result.
That's why you replicate.
As suggested, a good approach is to take p-values not as conclusive or decisive, but rather as a tool that must be supplemented by other statistics. In particular, the article emphasizes Bayesian methods, which can certainly provide additional information, but this approach can also be rather limited when priors are not well-defined or are entirely unknown, which is unfortunately often the case in many problem domains.
One potential question is how to determine the nature of the distinction mentioned in the conclusion between "preliminary research" and "confirmatory research", particularly in cases where statistics provide the primary evidence, as in, e.g. psychology. Further studies in the same vein as the preliminary research can certainly provide additional supporting statistical evidence, but this doesn't escape the problem that all of the evidence is probabilistic in nature. The key issue here is that since statistical approaches can only give probabilistic evidence that a hypothesis is correct, then they strictly cannot tell you what is certainly true, so even confirmatory research is quite open to falsification. So we wouldn't want the label of "confirmatory research" to somehow suggest to the public the idea that it is certainly correct.
In the presence of a poor prior the Bayesian probability would be biased in some way, so frequentists would say that the p-value in the absence of priors is actually superior in this case. Bayesians would reply that if they thought the prior might be poor then they would simply consider multiple different priors, but it's not clear how this would improve things much over the frequentist approach that simply assumes that the prior is unknown.
> So when you don't know the prior and you observe a low p-value on something, isn't that just "preliminary research" that needs to be further confirmed with other methods or at least the same test but using other data?
Yes, when you observe a p-value with low significance it should definitely indicate to you that more testing is necessary, either by using different testing methods, gathering new samples, or even just increasing the original sample size if that's possible. What I was trying to suggest in my last paragraph was that this should be the case even when we have highly significant p-values, because even significant p-values are not decisive. So even when we have "confirmatory research" that is highly statistically significant, we should still do all of the things that we would do when we have a p-value with low significance. It is sometimes the case that this subsequent research will overturn even very highly statistically significant results (though often this is unfortunately because mistakes in the original statistical methodology are uncovered).
The paper talks about how it seems researches are "hacking" (their word) p-values.If the researchers lack the ethics using one form of statistics, what is really stopping them from misusing Bayesian Analysis?
Sidenote: While we are talking about bayesian stuff. I recently ran into the sleeping beauty problem ( http://en.wikipedia.org/wiki/Sleeping_Beauty_problem ) and it showed how different interpretations of an experiment can lead to people believing in two very different answers in an almost religious way. Some people argued this thought experiment shows a clear flaw in bayesian thinking. I'm not sure either way, I'm still thinking about the puzzle.
I'd say I'm a halfer. I think the extreme sleeping beauty problem highlights how if you're woken up, it's not any more probable it's a head or a tails awakening, since, for a reason I can't explain, the million tails awakenings "don't accumulate".
In a sense, it's like the opposite of the Monty Hall problem, since here the sleeping beauty receives no information whatsoever during the experiment.
source : http://rfcwalters.blogspot.com/2014/08/the-sleeping-beauty-p...
<quote> BASP will require strong descriptive statistics, including effect sizes. We also encourage the presentation of frequency or distributional data when this is feasible. Finally, we encourage the use of larger sample sizes than is typical in much psychology research, because as the sample size increases, descriptive statistics become increasingly stable and sampling error is less of a problem. </quote>
In other words: report the effect size, plot the data, and increase the N.
Why Most Published Research Findings are False: http://journals.plos.org/plosmedicine/article?id=10.1371/jou...
Revised Standards for Statistical Evidence: http://www.pnas.org/content/110/48/19313.short
Similar thing to what happens in software security. Sure, you could tech people how to use memory management properly - and then watch them fail again and again. Or you could just provide an alternative solution like rust which changes the issue completely.
As I'm not a statistician by trade I don't keep up on the literature very well, but interconnectedness of data does seem to me to be a very important issue. I'm wondering if anyone can point me to some helpful reading to understand this side of the issue better. In particular, is there any approach to AB testing that can reliably address the issue of data interconnectedness in the kind of situation described above?
If one person visiting the site has no influence on other people visiting the site, then measurements of their behavior will be independent. If Facebook tests a different interface on half of their users and the changed behavior of those users indirectly has some impact on the behavior of the control group, then your measurements would have some level of dependence – I can imagine that this could happen but it's not clear to me that this would be a common scenario. The same would happen if you measure the behavior of the same person more than once – but in this case there's many procedures for working with paired or autocorrelated data.
Without better examples it's very hard to judge whether this is a real problem.
The example you give seems to me to oversimplify the issue of complex interconnections between data points, as if the traffic on a real website came from one set of referrals, while in reality its much more complex, with referrers inducing other referrers and a variety of campaigns, postings, etc. influencing each other, and over time, overlaid in a fairly complex pattern. In other words, a bunch of interrelated data, very little of which is actually independent of other items.
I'm not really asking for an explanation of this in the comment thread here; what I'd like to know is, if there are any studies or other publications that deal with the issue of how to evaluate tests run on interconnected data of this kind.
Also, statistics deals with many idealizations but the idea that randomization allows you to cleanly measure the effect of an intervention in the face of what would otherwise be confounding is simply not one of them. Sorry to disappoint, but with all you're telling us it simply sounds like the speaker was clueless.
I'm certainly not looking for "One Method for Interconnectedness Correction" (especially not, as you put it, with each word capitalized). I'm looking for studies or papers that might have addressed anything like the effect of interconnectedness of web data on AB testing. I think you're saying, you don't know of any, and also that you personally don't think it's a real issue.
Extreme example, but explains how the context can shift the findings significantly.
Interestingly enough there was a post on HN somewhat recently about Bayesian alternatives (BEST). The paper that was linked was: ftp://ftp.sunet.se/pub/lang/CRAN/web/packages/BEST/vignettes/BEST.pdf
And the recommended book I settled on was by the author of that paper (Doing Bayesian Data Analysis)
I feel like I'm "ahead of the curve" thanks to HN :D
Our prices are:
$10,000 random study, no results guaranteed. Not recommended! Highly likely to be damaging.
$20,000 basic study, p=0.5 - study inconclusive but implies it's at least not "more likely" that the damaging/negative result is correct. (Assuming uniform priors.) No scientific value.
$100,000 weak FUD. Suggests that the damaging/negative results (assuming bayesian reasoning or uniform priors) may be incorrect, at a suggestive p<0.10 level. Not conclusive and invites further studies which can strengthen damaging/negative result! Not recommended unless further studies are unlikely. Consists of ten studies, one of which is published with remaining buried. Unscientific.
$200,000 basic refutation. Refutes the damaging/negative result at a statistically significant p<0.05. Likely to be referenced and accepted. However, due to significance level, links to logicallee's getcher scientific results agency should still be avoided! (Invites skepticism.) Scientific. May take up to six months.
$500,000 Silver refutation. Highly significant refutation at p<0.02. Study can be extremely rigorous and sponsorship can be public. The results should be referred to and referenced as widely as possible. Unpublished results (the other 49 studies in the series) to remain unpublished and unreferenced. Due to the number of parallel studies to be involved, may take up to one year for these results. Highly scientific.
$1,000,000 Gold refutation. Highly significant refutation at p<0.01, with extreme amounts of data to be published. Should become the gold standard for data in the field, and the most significant effect published. Full refutation of damaging/negative result, with full scientific rigor. Should be widely promoted. May take up to 24 months to produce data. Gold standard of science.
ACADEMIC BONUS: prices are free for tenured professors, thousands of whom can do all the studies they want and only publish if they see some significant effect.
EDIT:
in other words, http://xkcd.com/882/