Strategies for making reproducible research the norm
elifesciences.org
elifesciences.org
Deincentivising an endless stream of trash publications will both cut down on the BS and narrow the field of research one might try to reproduce.
I don't have strong feelings about the optimal number of papers that should be published, but my sense is that we are running an order of magnitude above that number.
If a paper needed to clear the bars of
- can be reproduced by reading the paper and/or a few meetings with the original team
- data is independently reproducible to within the usual statistical parameters
It would presumably increase the signal:noise of research.
Requiring reproduction is several bridges to far.
But requiring reproducibility seems complete reasonable!
Even reproduction doesn't solve the problem of theoretical errors: https://news.ycombinator.com/item?id=36230450
I suggest an improvement.
What such a requirement would do is dramatically scale back research, and highly incentivize researchers to lie, or the "independent labs" to fudge in order to keep drawing funding from their collaborators.
I think a lot of money would be saved by avoiding researchers to waste resources in dead ends, by trying to build upon results that can't be reproduced. So not exactly double.
Bad incentive structures.
The point is, op is right that it’s expensive. The computer industry that we’re in claims to be data driven but I’ve observed numerous poor quality studies being done to drive decisions that I’m pretty jaded (no reproduction, poor sample sizes, skewed sample sizes where it’s employees, etc etc). And these are smart people where the decisions being made can impact the financial outcome.
But we don't even seem to be picking up the low hanging fruit - I used to work as an algorithm researcher, and reproduction there (for anyone who set up their experiments logically) is as easy as running a script and waiting x hours for the results to land. Yet reproduction studies were still a novel concept in that field, rather than a standard part of the submission workflow.
Furthermore, who's going to be doing that independent reproduction? How are you going to incentivize that?
If you require independent reproduction for every paper in a field where that's even possible, that's going to mean either you need to double the rate of running experiments in that field (unlikely), or double the number of researchers in that field (also unlikely) if you want to maintain anywhere close to the current rate of papers being published. (And again, even if you could double the number of researchers, you still need to provide enough of an incentive to reproduce that literally everyone is doing so for half of the studies they run. And you'd somehow have to ensure that you don't end up with fun, interesting studies getting 15 different groups reproducing them, while less-flashy studies that are still important fundamental science languish for years without anyone caring to try.)
Given how much of "publish or perish" is institutional inertia and culture, the idea that it could be systemically "nudged" to allow for this kind of shift is essentially a fantasy. Far too many of the Old Guard who are already in charge of the tenure committees would say, "Well, I had to publish 327 papers to be given tenure, and kids these days have so many newfangled things that let them write papers faster; why shouldn't we require them to publish 327 papers, too??"
You don't have to quite double the number of researchers or experiments. It typically takes less work to replicate final results than all of the work getting to the point of good results. Presumably a lot of the replication would be outsourced to CROs.
Though since the CROs would be focused on replication of results alone, any theoretical errors or errors in experiment design wouldn't be discovered. Replicating results will mostly catch fabricators. It won't catch nearly all of the far more important to the replication crisis bad experimental design.
It just seems that the entire scientific publishing industry is there to support a jobs and prestige program. Any science is just a side effect that somehow justifies the whole racket, from NSF budget to postdoc dinner table.
Societies fund many things for many reasons, not sure science is worth singling out here.
We could have a separate kind of journal for not-yet-reproduced results, and somehow ensure that the prestige is zero (or equivalent to just posting the study on a blog).
Suppressing publishing at the same time as actively trashing the ability to even study seems like a recipe for disaster way above publishing bad papers to me.
It's not impossible if we're actually interested in the truth.
> Suppressing publishing at the same time as actively trashing the ability to even study seems like a recipe for disaster way above publishing bad papers to me.
It's not clear that a study that cannot be replicated is worth the paper its printed on in the current environment where replication rates are 50% or less across the board. Other strategies that change how we approach individual studies are not as onerous (preregistration, open data), so maybe in that world individual studies would be worth it, but even in that world replication is the only sure fire way to validate results.
Replication is not a be all and end all, because things could replicate even though understanding is wrong (i.e., observation is right in the publication but everything else is wrong). Even the replication itself can be wrong, if certain materials share on unknown contaminant, for example. It addresses some issues, but not all.
For some areas it also not clear what replication means. What is it for theology, for example?
I can't imagine anyone who truly understands the current replication failure rates thinking that a single non-replicated study is valuable for anything other than informing what replications should be attempted.
Just think about it: replication rate is generally less than 50%. That means you'll have a better chance of determining the truth on any question posed by a study by flipping a coin than by actually reading the study.
> Replication is not a be all and end all, because things could replicate even though understanding is wrong (i.e., observation is right in the publication but everything else is wrong).
The observations are the only things that we have to get right. Interpretations can be subject to decades of debate (some QM debates are ongoing a century later), but if the data isn't reliable then you're just wasting time debating falsehoods. Replications are critical to ensuring we have reliable empirical data.
> Even the replication itself can be wrong, if certain materials share on unknown contaminant, for example.
Indeed, and literally the only way to figure out that such variables exist and are affecting the results are by independent replications. The more replications, the better. One replication is a bare minimum threshold to demonstrate that the process of gathering the data is at least repeatable, in principle.
> For some areas it also not clear what replication means. What is it for theology, for example?
It makes perfect sense in any context where you're gathering empirical data. For instance, if you're surveying people's interpretation of "free will" [1], then the process via which you probe their views should be replicable. This means being clear about the specific phrasing of the questions asked, the environment in which they were asked, the makeup of the cohort, and so on.
[1] https://www.researchgate.net/publication/274892120_Why_Compa...
Anything truly critical will (eventually) go through some replication/control of sorts (but it can take a long time).
You can either shut down most of research and then place your bets on what to keep and replicate, or you run broad but with a lot of incorrect stuff in it.
If you go for the former, you run the risk that you keep the wrong things, though. You have to have a way to quantify the direct and indirect costs of all the bad research and see if that trades off vs a much smaller research surface. Not sure if that is the case - empirical data matters much less for a lot of big decisions than people often make it out.
If it's inconsequential, then wouldn't that money be better spent on replications or other research that is consequential? I'm not really clear on what you're suggesting. Although maybe I wasn't really clear on what I've been suggesting.
Edit: to clarify, there are multiple ways to reorganize research. Consider an approach similar to physics, where there's an informal division between theoreticians and experimentalists. What if we have two different kinds of publications in social sciences, one that's proposing and/or refining experimental designs to correct possible sources of bias, and another type of publication that is publishing the results of conducting experiments that have been proposed. The experimentalists simply read proposals and apply for grants to conduct experiments, and multiple groups can do so completely independently. Conducting the experiment must strictly adhere to the proposed experimental design, no deviations can be permitted as is so common in social science when they find uninteresting results, otherwise this breaks the reliability of the results. A proposal should probably undergo a few rounds of refinement before experimentalists should feel confident in conducting the experiment, but I think the overall approach could work.
Sounds like a good idea to have a system of academic publishing that incentivises people to produce replications and similar studies then (and an academic norm of quoting multiple studies that support or oppose hypotheses)
And all that making any research that involves novel research unpublishable until someone else decides to dedicate their time to replicating the experiment from your little known working paper would achieve would be limiting incentives to experiment, especially in fields where it's perfectly possible to publish with statistical reexaminations of existing data (often flawed in other ways) instead.
I work in biology. At a panel of biology startup founders I heard one mention that she got a lot of her research ideas from papers studying bacteria which were published nearly a hundred years ago.
In biology you first seek to extend published results. Only if the extension attempt fails would you spend effort trying to replicate it (assuming you just don't abandon the pathway entirely).
And this is in a field where everything is based on code, where in principle reproducibility is easy. Go into materials science or chemistry and try to synthesize something following a published paper and you get all sorts of problems. Different equipment, different temperature, not all steps documented, ... Reproducing experimental findings can take you months.
So if you feel some impactful work is suspicious .. I think disproving it would absolutely be incentivised
If you show its actually correct.. Well then usually it's not that hard to push the envelope a bit further and say something new. That happens all the time
Also, the LK-99 example is an exception, not the norm–the chances of receiving significant attention for a replication study are near zero in almost all other cases.
Even in the ideal world, you effectively almost never end up with a replication paper. Either it replicates and you add on your own novel research. Or it doesn't replicate and you discover something new
You can in theory end up with a super dull null result that disproves someone else's claim. But even then, when you set out on the project you're aim is to add something new on top of what's been already done. This happens all the time
About two hours is the cumulative time one must cater to the Dockerfile for a 3 weeks project.
But it requires institution insisting on reproducibility, and fostering best practices to make it even easier for the researchers to be compliant.
I get it that reproducibility can be quite hard for biology. But ML cannot be taken as an example of a hard problem.
The problem is knowing upfront which of your work would need to be reproducible, or having the discipline to do all your hacking starting from such reproducible setups.
Do you know much about how reproducibility is approached in Julia? Maybe hold off on calling it a lie if you're not experienced in what you're talking about.
Scientific reproducibility requires only that versioned binaries be functionally equivalent if they have the same version, which is quite independent of this and certainly exists in Julia.
Would love a link to the Guix mailing list discussion, if you can dig it up.
On the other hand, it really is the case that there's just not much of any value in learning that [shocking possibility] is, as everybody would naturally expect, indeed not the case. And filling up limited journal space with such discoveries would seem to be counter-productive, at best. And when you have limited space/funding for researchers, one guy who keeps proving everything everybody knows to be false, to be false, is always going to be perceived as less valuable than one making [shocking discovery] [... which ends up being proven false years later].
The repetition team being incompetent sounds like a cop out. The researcher did a bad job and it’s on them to explain better etc in that case. No excuses, if it can’t be reproduced it isn’t taken seriously no exceptions
Furthermore, I think you can often see poorly done science in the papers themselves. They will use suggestive wording in surveys, unreliable sources for sampling such as Amazon Mechanical Turk, and maybe one of the biggest tells is measuring a large number of unnecessary variables. That does very little to further your experiment, but absolutely ensures you can p-hack your way to a statistically significant result. Another is ignoring such patently obvious viable confounding issues, that one can't reasonably appeal to Hanlon's razor.
[1] - https://en.wikipedia.org/wiki/Replication_crisis#In_psycholo...
[2] - https://www.google.com/search?q=site%3Anytimes.com%20%22Jour...
But this simply isn't true in physics where negative results are very common. This is at least an existence proof that this can work, people just have to get their heads straight on what research means.
So, for example, suppose negative results become as valuable: well, they are easier to produce. They are also less valuable as stepping stones for further research. Given that, you'd still need to have a metric that compares publishing positive results to negative results. Even if you declare them to be equally important, the shared understanding will be that they aren't. And one would be more important than the other. And here were are back to square one.
There are some minor things that can be done in the near future. For example, results produced with code must come with the code that produced these results. A lot of research bodies resist this because they want to commercialize their code, or their code may inadvertently contain organization's secrets and therefore needs more auditing... but, in the end of the day, it needs to be made clear that this is a necessary and unavoidable price to pay.
Data sharing is even more problematic. Beside confidentiality concerns, data is always a bargaining chip in the game of getting collaborators (and grants). Should it be made public, it loses its value to those who collected it. Right now, the trend is: if you managed to collect a worthwhile dataset, then you'll cover yourself foot to head with NDAs, contracts of all kinds etc, and will sit on it, exploiting it for a series of research. And if anyone wants to do research on the same subject, you will only invite them if they bring grants or equipment etc.
But you cannot really verify results w/o having the data available. Even if you have the code.
---
It's really sad to see how research is doing wrt' programming in part because of the above, but I don't think the programs outlined in OP will have a noticeable effect. They don't paint a convincing picture in terms of incentives, i.e. they don't answer the question why would researches want to do any of that RepRes and OS training. Even in computationally-heavy research today you often find that all the computation work is outsourced by the researchers and they themselves have no clue what their code is doing.
Above were all sorts of arguments for why the current (or yours) approaches are ineffective. But I don't claim to know what needs to be done.
We don't need to be threatened by the mere existence of fake journals that publish articles with "counterfeit consciousness" in them. No one actually reads that stuff. It only exists to feed badly-managed communities with terrible incentives. And the solution to such bad communities is to create good ones and showcase their work. Not waste our time obsessing over the fear that somewhere, someone is not a great researcher.
Poor research is the norm and should be ignored. Good research is the exception that should be recognized and nurtured.
If it becomes a habit, and a natural expectation, then teaching the specifics pertaining to any specific field at an advanced level should be easier. And maybe it would encourage the public to demand better science reporting.
Also, mark the topics in the introductory textbooks that are not based on reproducible research, so that students know the actual state of the field that they're studying.
I do industrial R&D, and my stuff doesn't get published at all, but I benefit from doing work that is at least "open and reproducible" within my organization. It actually improves the quality of my work.
The reason why a lot of people don't make reproducible research is the same reason why communism tends not to work.
In order to make it work you need to put incentives in place. For example. No journal publishes a paper unless the work has been reproduced by someone from a different institution.
It generates any data plots/tables through CI, which means you need to have your data available and consumable, and containerised code. Its PR reviewers are scientists doing peer review.
Then you get public review and comment, reproducible environments, and open data.
JOSS has a review process like this, but I don’t think I’ve heard of any journals focused on primary scientific results using this kind of GitHub based review process
https://joss.readthedocs.io/en/latest/submitting.html#the-re...
I wonder if it could also get people to sponsor replication studies, and publish them linked alongside.
The problem with academic research into predicting speeds is not as much the lack of reproducibility, but the lack of measures that make sense.
Hundreds of papers compare a short-term prediction (say 30 minutes ahead) using mean average error, or similar measures. This completely ignores the fact that vehicle speeds are mostly constant and only vary significantly during rush hours or other disruptions, at which time the process of breakdown is rather chaotic.
From what I gather this will not change easily, because it is considerable academic fun to improve on a number for existing benchmarks, instead of redefining the game.
I'd love to help academics in this field who need some guidance on practical applications!
Well that's confusing, because reproducibility is a core tenet of science. How else could you tell whether your results were a statistical fluke or due to some flaw in the study design?
Why not require it?