Think about all of the fancy dependencies a piece of code might need to run. Think of all of the processing steps some data needed to go through. Think of the work required to set up some computational cluster without which the work just doesn't happen. Doing it once is hard enough, documenting it so anyone else can do it is maybe an order of magnitude harder.
Which is a huge part of why the replication crisis is such a thing (besides the outright fraud and the publication bias). The very fact that the datasets and codebases are so disgusting is precisely why the results coming from that data can't be trusted.
Is it though? From what I’ve seen it is mostly caused by p-hacking/small sample sizes/poor experimental design.
An incomplete list of other causes might contain:
• Wrong data or maths. A remarkably large number of papers (e.g. in psychology) contain statistical aggregates that are mathematically impossible given the study design, like means which can't be calculated from any allowable combination of inputs. Some contain figures that are mathematically possible but in reality totally implausible [1]
• Fraud. I've seen estimates that maybe 20% of all published clinical trials haven't actually been done at all. Anyone who tried to figure out the truth about COVID+ivermectin got a taste of this because a staggering quantity of studies turned out to exhibit disturbing signs of trial fraud. Researchers will happily include obvious Photoshops in their papers and journals will do their best to ignore reports about it [2]
• Bugs. Code doesn't get peer reviewed, sometimes not released either. The famous Report 9 Imperial College London COVID model was in development since 2004 but was riddled with severe bugs like buffer overflows, race conditions, and even a typo in the constants for their hand-rolled PRNG [3]. As a consequence their model produced very different numbers every time you ran it, despite passing fixed PRNG seeds in on the command line. The authors didn't care about this because they'd convinced themselves it wasn't actually a problem (and if you're about to argue with me on this, please don't, if I have to listen to one more academic explaining that scientists don't have to write deterministic "codes" I'll probably puke).
• Pretend data sharing. Twitter bot papers have perfected the art of releasing non-replicable data analysis, because they select a bunch of tweets on topics Twitter is likely to ban, label them "misinformation" and then in their publicly shared data include only the tweet ID, not the content. Anyone attempting to double check their data will discover that almost all the tweets are no longer available, so their classifications can't be disputed. They get to claim their analysis is replicable because they shared their data even though it's not.
• Methods that are described too vaguely to ever replicate.
And so on and so forth. There are an unlimited number of ways to make a paper that looks superficially scientific, but doesn't actually tell us something concrete that can survive being double checked.
[1] e.g. https://hackernoon.com/introducing-sprite-and-the-case-of-th...
[2] https://blog.plan99.net/fake-science-part-i-7e9764571422
[3] https://dailysceptic.org/archive/second-analysis-of-ferguson...
I would wager that the vast majority of analyses / code, in most fields of research, can run on a basic laptop. If I had to put a number on it, I'd say >99.999% of papers.
Not in terms of proportion of data of course, but in terms of proportion of papers.
And I was the most organized programmer among my peers by a long shot. Judging from the messiness of shared code, I can’t imagine how bad other people’s unshared work is.
Some of it is just really clean Python code that generates a bunch of simulated data to train some machine learning code on with human readable configuration files.
But neither one of them involve the words "$COUNTRY Ministry of Health approval...", which kicks things to a whole new level.
Maybe the goal should merely be for the author to be able replicate the results from the torrent archive (for a fee of course). That way anyone who needs the conclusion to be correct can buy the validation.
1) Are you sure the data is deidentified? Can your analysis be done on the deidentified data?
2) Do you actually own the data? This is more complicated than "Well you published a paper on it...". The data is likely governed by data use agreements. Those are often extremely complex - especially at the international level. Or if you're working with particular populations - they're not wrong to be wary of who their data is given to, given the history of how science generally has treated them.
3) Is the data in a form that's genuinely suited to being shared? Is it in a flat file, with clear and easy to understand variable names, a good data dictionary, etc.? Is missingness, and the reasons why, clear and evident (the difference between a variable with a lot of missingness and a variable that's a bit garbage is often subtle)?
4) Where do you put it - and importantly, who pays for that, and how long?
I agree generally with data and code sharing policies (all of my recent papers include zenodo deposits of data, and I work on a large open source simulation code). However, there are some serious issues here:
1. Code licensing is not always clear. Sometimes you are working with a code that some PhD student wrote, who left the field, whose advisor shared it with you. Can you share this code? Not necessarily... but this problem will go away with time.
2. Some codes include proprietary components, and chopping them out can make the code inoperable (or useless). One notable instance of this is a very famous hydro code that was used to produce predictions for nukes, that was then repurposed for running simulations of planetary collisions.
3. Sharing data is prohibitive, and if you don't have the data having the code might not really help. Our projects can produce Pbs of data, costing 10s-100s Ms of CPU hours, and as such making the data open source can be really tricky. This then means you can only share a selection of runs. On the other hand, even if you had the code, you can't reproduce a lot of the results, because of the cost of running the simulations... This actually will get _worse_ over time as we run bigger and bigger simulations.
Isn't it the university who owns everything? I don't know about astrophysics, but it's a big problem in engineering. Universities have been known to patent and resell research made by msc/phd student and profit from it without paying any compensation.
Consider this: student at institution A (UK) writes code, postdoc at institution B (USA) modifies it, shared with researcher at institution C (China) who runs it for a paper first-authored by student at institution D (Germany).
Who has the rights to share the code and data? Which jurisdiction would this fall under?
Because the incentive in academia is to take a single idea that kinda works and parcel it out into as many, generally low quality, published papers as you can.
The moment you release the data somebody else can start parceling out those papers instead of you.
- What benefit does this give to the researchers who are publishing the data?
- Who is paying for storing the data - frequently in the TB?
When the answers are "approximately none", and "the researchers" you're not going to get many takers.
And that's assuming it's easy, it reality it often won't be. If there's a lot of data, it's literally just a pain to upload it (network bandwidth). If there's sensitive data (PII), it's a pain to redact it and make sure you aren't leaking any. Data is frequently in strange formats, and it's a pain to translate it to a standard one. Etc.
---
I've worked with 3 university labs as a contract programmer. In all of them I worked with data with one of the issues mentioned above. Health information in one, TBs of photonics data in another (which was being parsed by extremely janky code too), and 4 16-bit channel images in the last (hundreds of GB of them too). Admittedly for the last it would have been easy to upload it as long as you didn't want people to be able to actually view it (on the other hand I wrote some software to let people false color them live in a browser for the lab, so that the researchers could view them).
Which answers the "what benefit does it provide to the researchers publishing the data" question. A quick search answers the funding question as well, it's funded by the NIH, not the individual labs using it.
I think this example supports my point. The NIH came up with a way to give different answers to the two questions I asked, and it gets used. I'm glad the NIH has been making this a thing, it's a great use of public funds.
I'd still caution anyone from trying a "make a data platform and researchers will use it" approach to the problem unless they can answer those questions.
I'm not sure how this could work with things like photos etc (though there are plugins for some of this)