Meaning what?
> I'm imagining something like you have a dataset and you have to upload that dataset to some third party that checks it for it's validity.
The results are what they are. What is 'validity' meant to mean?
> Of course, this is a completely silly idea but I'd love to know if someone has like any tangential related thoughts on this
Quantitative studies are already published with the proper analyses, which are invariably produced 'automatically' using software, not manual methods.
I imagine there might be some value in publishing raw data, though. There may sometimes be questions like privacy, but I don't imagine they'll always be show-stoppers.
> Then you write exactly WHAT you are going to do with the data and then you send the proposed "routines"
this roughly matches up with the idea of "preregistration" of research, i.e. you define and share what your experimental method is going to be before you start looking at the data and performing analysis, to help guard against some unconscious or conscious decisions to adjust the method after you have observed the experimental data.
i've never heard of cos.io before but they're a top hit for me when searching for "preregistration" https://cos.io/prereg/
Andrew Gelman has written quite a lot about related topics in the past (the "replication crisis" - especially in the social sciences, "p hacking", "the garden of forking paths" http://www.stat.columbia.edu/~gelman/research/unpublished/p_... , https://andrewgelman.com/2017/03/09/preregistration-like-ran... )
I quite like how Gelman theoretically frames this in his "forking paths" paper:
Consider the following testing procedures:
1. Simple classical test based on a unique test
statistic, T, which when applied to the observed
data [ y ] yields T(y).
2. Classical test pre-chosen from a set of
possible tests: thus, T(y; phi), with
preregistered phi. For example, phi might
correspond to choices of control variables in a
regression, transformations, and data coding and
excluding rules, as well as the decision of
which main effect or interaction to focus on.
3. Researcher degrees of freedom without fishing:
computing a single test based on the data, but in
an environment where a different test would have
been performed given different data; thus
T(y; phi(y)), where the function phi(.) is
observed in the observed case.
4. "Fishing": computing T(y;phi_j) for j=1,...,J:
that is, performing J tests and then reporting
the best results given the data, thus
T(y; phi^{best}(y)).It's absurd that reproducing results be any more of a challenge than simply re-running a freely-available program.
Software is unique in that it can be generally be duplicated and executed trivially.
There's no way to make it trivial to reproduce a test on the strength properties of a new ceramic. There is a way to do this for software, and it's rather silly that it isn't standard scientific practice to do so.
I realise I'm taking a strong line here, but I've never seen a good argument against it.