On the other hand, it really is the case that there's just not much of any value in learning that [shocking possibility] is, as everybody would naturally expect, indeed not the case. And filling up limited journal space with such discoveries would seem to be counter-productive, at best. And when you have limited space/funding for researchers, one guy who keeps proving everything everybody knows to be false, to be false, is always going to be perceived as less valuable than one making [shocking discovery] [... which ends up being proven false years later].
But this simply isn't true in physics where negative results are very common. This is at least an existence proof that this can work, people just have to get their heads straight on what research means.
Furthermore, I think you can often see poorly done science in the papers themselves. They will use suggestive wording in surveys, unreliable sources for sampling such as Amazon Mechanical Turk, and maybe one of the biggest tells is measuring a large number of unnecessary variables. That does very little to further your experiment, but absolutely ensures you can p-hack your way to a statistically significant result. Another is ignoring such patently obvious viable confounding issues, that one can't reasonably appeal to Hanlon's razor.
[1] - https://en.wikipedia.org/wiki/Replication_crisis#In_psycholo...
[2] - https://www.google.com/search?q=site%3Anytimes.com%20%22Jour...
The repetition team being incompetent sounds like a cop out. The researcher did a bad job and it’s on them to explain better etc in that case. No excuses, if it can’t be reproduced it isn’t taken seriously no exceptions
And this is in a field where everything is based on code, where in principle reproducibility is easy. Go into materials science or chemistry and try to synthesize something following a published paper and you get all sorts of problems. Different equipment, different temperature, not all steps documented, ... Reproducing experimental findings can take you months.
The problem is knowing upfront which of your work would need to be reproducible, or having the discipline to do all your hacking starting from such reproducible setups.
Do you know much about how reproducibility is approached in Julia? Maybe hold off on calling it a lie if you're not experienced in what you're talking about.
Scientific reproducibility requires only that versioned binaries be functionally equivalent if they have the same version, which is quite independent of this and certainly exists in Julia.
Would love a link to the Guix mailing list discussion, if you can dig it up.
About two hours is the cumulative time one must cater to the Dockerfile for a 3 weeks project.
But it requires institution insisting on reproducibility, and fostering best practices to make it even easier for the researchers to be compliant.
I get it that reproducibility can be quite hard for biology. But ML cannot be taken as an example of a hard problem.
So if you feel some impactful work is suspicious .. I think disproving it would absolutely be incentivised
If you show its actually correct.. Well then usually it's not that hard to push the envelope a bit further and say something new. That happens all the time
Also, the LK-99 example is an exception, not the norm–the chances of receiving significant attention for a replication study are near zero in almost all other cases.
Even in the ideal world, you effectively almost never end up with a replication paper. Either it replicates and you add on your own novel research. Or it doesn't replicate and you discover something new
You can in theory end up with a super dull null result that disproves someone else's claim. But even then, when you set out on the project you're aim is to add something new on top of what's been already done. This happens all the time
So, for example, suppose negative results become as valuable: well, they are easier to produce. They are also less valuable as stepping stones for further research. Given that, you'd still need to have a metric that compares publishing positive results to negative results. Even if you declare them to be equally important, the shared understanding will be that they aren't. And one would be more important than the other. And here were are back to square one.
There are some minor things that can be done in the near future. For example, results produced with code must come with the code that produced these results. A lot of research bodies resist this because they want to commercialize their code, or their code may inadvertently contain organization's secrets and therefore needs more auditing... but, in the end of the day, it needs to be made clear that this is a necessary and unavoidable price to pay.
Data sharing is even more problematic. Beside confidentiality concerns, data is always a bargaining chip in the game of getting collaborators (and grants). Should it be made public, it loses its value to those who collected it. Right now, the trend is: if you managed to collect a worthwhile dataset, then you'll cover yourself foot to head with NDAs, contracts of all kinds etc, and will sit on it, exploiting it for a series of research. And if anyone wants to do research on the same subject, you will only invite them if they bring grants or equipment etc.
But you cannot really verify results w/o having the data available. Even if you have the code.
---
It's really sad to see how research is doing wrt' programming in part because of the above, but I don't think the programs outlined in OP will have a noticeable effect. They don't paint a convincing picture in terms of incentives, i.e. they don't answer the question why would researches want to do any of that RepRes and OS training. Even in computationally-heavy research today you often find that all the computation work is outsourced by the researchers and they themselves have no clue what their code is doing.
Above were all sorts of arguments for why the current (or yours) approaches are ineffective. But I don't claim to know what needs to be done.