And study preregistration to avoid p-hacking and incentivize publishing negative results. And full availability of data, aka "open science".
And study preregistration to avoid p-hacking and incentivize publishing negative results. And full availability of data, aka "open science".
I do agree though, negatives are just as important when the intent is to prove/disprove a meaningful hypothosis.
we tried using 0.11 mL, it didn't work
we tried using 0.12 mL, it didn't work
we tried using 0.13 mL, it didn't work
we tried using 0.10 mL, it didn't work
we tried using 0.11 mL, it didn't work
we tried using 0.13 mL, it didn't work
we tried using 0.15 mL, it didn't work
we tried using 0.17 mL, it didn't work
we tried using 0.16 mL, it didn't work
we tried using 0.18 mL, it didn't work
we tried using 0.20 mL, it didn't work
we tried using 0.14 mL, it didn't work
we tried using 0.12 mL, it worked so we published
Do you want to know the ones that "didn't work" existed? Or are you happy with just the one that "worked" being written up in isolation?E.g. if you search for eggs and cholesterol you should find all studies with their summarized results on whether eggs are ok or not for your cholesterol, grouped by researcher, so if somebody does 200 studies to find the one positive it's instantly visible.
Improving the quality of measurements and data could be a rewarding pursuit, and could encourage the development of better experimental technique. And a good data set, even if it doesn't lead to an immediate result, might be useful in the future when combined with data that looks at a problem from another angle.
Granted, this is a little bit self serving: I opted out of an academic career, partially because I had no good research ideas. But I love creating experiments and generating data! Fortunately I found a niche at a company that makes measurement equipment. I deal with the quality of data, and the problem of replication, all day every day.
One could make the case that in GWAS studies it has occured, but not because small effect sizes are inconsequential, the statistical methods just weren't able to separate grain from chaff for a while.
An allele that is responsible for 2% of the variation in disease risk might seem inconsequential, but 25 of those together can serve as a polygenic risk score that can predict disease and target treatment.
Of course they're stupid. Everyone is stupid. That's why we have a "scientific method" and a formal discipline of logic to overcome fallacious reasoning and cognitive biases. If people weren't stupid we wouldn't need any of these disciplines to check our mistakes.
And yes, what you describe does happen all of the time. We literally just had a thread on HN about the failure of the amyloid hypothesis in Alzheimer's and the decades of work put wasted on it. Many researchers are still trying to push it as a legitimate therapeutic target despite every clinical trial to date failing spectacularly. As Planck said, science advances on funeral at a time.
Which isn't to say that small effect sizes aren't legitimate research targets either, but if you're after a a small effect size, the rigour should be scaled proportionally.
https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3444174/
Quote:
> A commonly cited example of this problem is the Physicians Health Study of aspirin to prevent myocardial infarction (MI).4 In more than 22 000 subjects over an average of 5 years, aspirin was associated with a reduction in MI (although not in overall cardiovascular mortality) that was highly statistically significant: P < .00001. The study was terminated early due to the conclusive evidence, and aspirin was recommended for general prevention. However, the effect size was very small: a risk difference of 0.77% with r2 = .001—an extremely small effect size. As a result of that study, many people were advised to take aspirin who would not experience benefit yet were also at risk for adverse effects. Further studies found even smaller effects, and the recommendation to use aspirin has since been modified.
Long-term aspirin use has its own risks, like GI bleeds, and the MI benefits are clearly not warranted given those risks.
> There was a 44 percent reduction in the risk of myocardial infarction (relative risk, 0.56; 95 percent confidence interval, 0.45 to 0.70; P<0.00001) in the aspirin group (254.8 per 100,000 per year as compared with 439.7 in the placebo group).
I agree if you said from the start you meant general incentives, especially in pharma development, but that is by and large a different conversation.