'Generous’ approach to replication confirms many social science findings
sciencemag.org
sciencemag.org
I'd sure say so. Statistics in science is just so broken. Lots of inappropriate techniques chosen on tradition, ignorance (or concealment) of multiple comparisons, the base-rate fallacy, etc, etc. It's not just social sciences, but medicine, biology, and just about every field that relies greatly on lower-power studies.
Bayesian techniques are really called for in those situations.
What it means is that if I trust a study - any study (from my perspective) - I'm essentially flipping a coin. If non-scientist citizens can't rely on what is coming out of the field, then it seems like a massive problem that needs solving before any other.
That's why have meta-analysis and systematic reviews and even those aren't exactly bullet-proof.
I work in fluid dynamics. In one subfield I work in there is at least one confounder in most studies (Weber number and Reynolds number are confounded, in particular) and an important variable is often omitted (turbulence intensity). Many people don't seem aware that turbulence intensity matters, either, and most who do think it matters don't want to measure or estimate it, so it tends to be ignored. Just because it's hard to measure does not mean it's not important! This is supposed to be "hard" science, but sometimes I feel it's not much better than social science.
How long has this been a problem? I'd say over 50 years, easily!
At a recent conference, one of my papers addressed these issues briefly (I tried to avoid them in a data analysis), but I think it'll take more than that to correct these problems. So I'm planning a new article specifically addressing these concerns. I still expect progress to be slow, but any movement in the right direction would be appreciated...
You should be skeptical of any so-called science that cannot actually rigorously apply the scientific method, be it for practical, ethical, or any other reasons.
Sadly, I think there is huge incentive to maintain status quo. It is essentially a loophole to manipulate the scientific consensus without the actual scientists being dishonest. Some wise once said "When a measure becomes a target, it ceases to be a good measure.". I think that what we call as "science" today, has thus that fallen prey to this.
Medecine is more problematic. Many studies are at the border of statistical significance but it is impossible to conduct large scale experiments with humans, so it’s not like if we have a better alternative.
You're being downvoted, but my understanding is that lots of social science researchers openly refer to themselves and their discipline as activism. Am I mistaken? Is this observation controversial?
EDIT: Now that people are downvoting me, perhaps someone could take a moment to explain where I'm mistaken?
That's a pretty strong contrast to say "We have found that X is harmful, and will advocate for people having reduced exposure to X."
Consider, for example, the number of infectious disease epidemiologists, vaccine developers and doctors who have essentially been forced to engage in "activism" thanks to anti-vaxx movements.
That's a troubling suggestion. Results are valuable to the extent that they're both accurate and surprising. To systematically suppress surprising results as a negative predictor of accuracy sounds like a formula for suppressing surprisingly valuable papers.
Andrew Gelman talks a lot about this issue on his blog:
https://andrewgelman.com/2014/08/01/scientific-surprise-two-...
"Weeding out" doesn't mean systematically suppressing. It simply means exhibiting caution before putting the paper in a reputable journal. The Internet provides the entire system with an escape valve in scientists' capacities to publish papers on their own websites. If the finding is surprising enough, it shouldn't be problematic attracting some attention, particularly in the social sciences.
It does when I'm weeding my garden. A result that is weeded out of a journal is suppressed from taking root in the minds of its readers. I agree that extraordinary claims require extraordinary evidence to accept. But as long as they meet the ordinary standard of rigor it's just those claims that inquiring readers want to entertain.
This is an obvious way to tamper with the results; it's just more of the same kind of p-hacking that bad researchers are so often doing. They are using "we re-do the study with a larger population" as a way to re-roll the dice if the first die roll doesn't come up the way they want. (Note that if the die roll did come up the way they want, they don't re-do the study with a larger population in order to see if the replication fails).
Nobody should be taking this seriously.
To give a simple model, suppose you decide an effect is significant if there's only a 5% chance that you'd see this data if the effect didn't exist. If you run the experiment again with the same threshold for significance when don't get the desired result, then the probability of seeing an effect that doesn't exist rises to 9.75% (= 0.05 + 0.95 * 0.05).
The effect isn't merely "not 100% unproblematic" it's a serious problem! You've gone from what looks like a p-value of 5% to a p-value of 9.75%.
The fact that the second study is done with a larger sample is pretty much irrelevant unless it also comes with a higher p-value threshold for you to accept the result.
Is this common in modern replication studies? Do the results of the prediction-market ever get announced/published prior to the replication results themselves?
> Both approaches did well at predicting the outcome for individual studies, and they predicted an overall replication rate very close to the actual figure of 62%.
I think the real headline here is that we should supplement peer review with replication prediction markets.
Science, as a process, naturally replicates studies that are worth replicating, and drops those that are not. The idea of "prediction markets" is a new-tech way of doing what scientists already do: decide where to invest their (limited) budgets to maximize long-term career success.
Edit: got the title wrong
* Prior replication studies replicated 39% of papers in psychology journals [Ed note: That doesn't mean the other 61% were complete failures; most just didn't produce results as statistically strong as the originals IIRC] and 61% in economics journals.
* This replication study greatly increased the number of participants in some experiments, and for two papers, that changed the replication results from failure to success. Overall 62% of the studies were replicated successfully.
* One researcher "points out that the project repeated only one experiment from each paper, and in his case, it wasn’t the strongest or the most important."
* If an initial replication attempt failed, the researchers added even more participants.
Start with 25 participants, then if you don't get the result you want, add 50 and try again. Carry on doubling until you get the result you want.
Not a statistician here, but as each addition is essentially a new trial, shouldn't they have applied Bonferroni correction to the results?
For any arbitrary paper, what is your assumption about its accuracy? How much do you rely on it? Can you put a number to it? The null hypothesis that research papers, especially in very difficult fields like social science, are near-infallible seems to be the error, AFAICT. It seems like something people outside science would assume, but I see scientists who are surprised by a 62% replication rate. I'm not surprised or concerned, but maybe I should be.
Why do they call increasing the number of participants, "generous"? Doesn't that increase accuracy, and isn't accuracy the whole point of the replication study? Generosity implies some sort of favor, something above and beyond, while in this case it seems necessary based on the fact that the results changed - it would be failure of the replication study to not increase the number of participants.
There is no such thing as "an arbitrary paper". It will depend very much on the study in question - including soft factors like who wrote it, but also the study design, sample size, if I think they approached the statistics appropriately, etc.
"Why do they call increasing the number of participants, "generous"? Doesn't that increase accuracy, and isn't accuracy the whole point of the replication study? Generosity implies some sort of favor, something above and beyond, while in this case it seems necessary based on the fact that the results changed - it would be failure of the replication study to not increase the number of participants."
The reason it might be thought of as generous is the criteria they determine for saying something is replicated is a significant effect in the same direction as the original study (for the record, I hate this criterion). A larger study is more likely to find a significant result if indeed there is one there, so they're giving studies that report an effect a very strong chance of seeing that replicated if indeed it was real.
From what I've seen in academia, it varies depending on a number of variables, like who wrote the paper and whether their results help or hurt my research.
Scientists are human, after all. And in my discipline, it was rare that the paper would have enough concrete details to replicate, so this way of thinking was the norm.
The unstated implication here is that it seems a very large chunk of social science researchers are faking their numbers, likely through p-hacking, to create attention (and journal) grabbing headlines that are, for lack of a better word, simply fake.
Knowing that I have no strong reason to trust most of the conclusions is useful. But know which papers I can trust is more useful.
Anyone knows of such a list?