Preregistration of clinical trials causes medicines to stop working
chrisblattman.com
chrisblattman.com
* In 2000, a rule was passed that clinical trials must include their hypotheses (i.e. clinical outcomes) in their initial registration before the trial is conducted.
* From the graph, the number and proportion of clinical trials through NLHBI with positive outcomes (i.e. an improvement in some health condition) is far lower after 2000 than before.
It was previously easy for researchers to conduct a trial, and then find positive trends in the data to form positive outcomes by framing the data from the trial in a different manner, as if the trial were originally intended to test that positive outcome. This is a problem related to A Priori Probability, where you're essentially turning research, which could have been statistically applied to a larger population given some statistical confidence, into deductive reasoning which doesn't necessarily apply beyond the finite sample size [1] (also see @entee's comment and links for better explanation of why this is important).
This article is stating that the graph implies that, because there are fewer positive outcomes after this new rule in 2000 than before, most of the positive outcomes prior to 2000 (which likely lead to new FDA-approved drugs on the market) must have been illegitimate.
I would point out though, that at least from that graph, the last positive outcome shown prior to 2000 was in 1996, so it's not like it obviously continued right up until 2000 and then suddenly stopped. Even when that rule was imposed in 2000 (I'm not necessarily sure how the ruling was phased in), there was a 4-year gap since the last positive outcome before then, and I don't see another 4-year gap prior to that since before around 1979. I'm not saying I disagree with the article (obviously), but just that while the graphic used seems to imply a trend, it seems like more data would be needed to conclusively show that trend.
From the abstract:
> We identified all large NHLBI supported RCTs between 1970 and 2012 evaluating drugs or dietary supplements for the treatment or prevention of cardiovascular disease.
So this is for a narrow field where there are effective drugs already and where all the low hangoing fruits are fully picked because the business opportunity is huge.
And you can tell that by looking at the graphic...? There's also a reference to a X2-test with p=0.0005 (meaning that there's a 0.05% chance that the trend you see is accidental and would disappear with more data).
I think you're right to be skeptical (correlation is not causation and so on...), but not that too little data is the problem here.
Maybe you are using these terms differently than they are used professionally, but this is not correct if it is meant as a criticism of Bayesian methods.
There is nothing inherently better (and in fact there are several things potentially worse) about looking at the data with inductive reasoning about a potentially larger population with assumed baseline properties.
The problem here is simply p-hacking, whether by accident or intentional. This can be done whether you are performing Bayesian analysis involving a prior distribution, or performing frequentist analysis with assumptions about null hypothesis and subjectively chosen rejection criteria.
If you don't register your model validation procedures ahead of time, then after the fact you can data mine among many (or even all possible) model validation scores and only publish scores that show a positive effect.
It's not about one framework of statistics or the other. It's about statistical rigor for whatever framework. It's about pursuing more than one study. And it's about attempting to use holistic model validation techniques that cover wide ranges of anticipated outcomes, to hopefully drive down the chances that you develop a conclusion from the study based only upon a narrow set of outcomes that could be biased.
It was an attempt to explain in more lay terms how when p-hacking is performed on some dataset, the "findings" can rarely be applied to larger (or different) datasets due to the arbitrary criteria used to reject the parts of the data that aren't conducive to the selected outcome, since the rejection criteria were sort of custom-fit to that specific dataset.
I agree with everything you said.
See these (kind of old but still quite appropriate) links:
http://journals.plos.org/plosmedicine/article?id=10.1371/jou...
http://www.theatlantic.com/magazine/archive/2010/11/lies-dam...
It's this kind of work that led to pressure to pre-declare statistical objectives, and now we see the results. Biology is hard, medicine is hard, we must be quite humble about what we understand in these fields and therefore how easy it is to be misled by early promising results.
The punchline is: once the researchers had to declare their statistical procedure before turning in the results most of the dots get colored "null." Meaning the strongest effect in the past was the ability to argue your way into a more favorable analysis.
I really hope the article is satire.
http://www.nature.com/news/over-half-of-psychology-studies-f...