Study without original raw data or source code, is just authors opinion, not a science!
Study without original raw data or source code, is just authors opinion, not a science!
For my thesis, I measured the Paterson function for a series of colloids. I can imagine other scientists finding this useful and I'd be happy to submit it. However, it's not the raw data. What I actually measured is the polarization of a neutron beam, which I then mathematically converted into the Patterson function. So I should probably submit the neutron polarization I measured, so that other scientists can check my transformation. Except that I can't directly measure the polarization - all I really measure are neutron counts versus wavelength for two different spin states, so that must be my raw data. But those counts versus wavelengths are really a histogram of time coded neutron events. And those time coded neutron events are really just voltage spikes out of a signal amplifier and a high speed clock.
If a colleague sent me her voltage spikes, I'd I'd assume she was an idiot and never talk to her again. Yet, I've also see experiments fail because of problems on each of these abstraction layers. The discriminator windows were set improperly, so the voltage spikes didn't correspond to real neutron events. The detector's position had changed, so the time coded neutron events didn't correspond to the neutron wavelengths in the histogram. A magnetic field was pointed in the wrong direction, so the neutron histograms didn't give the real polarization. There was a flaw in the polarization analyzer, so the neutron polarization didn't give the true Patterson function. And all of this is assuming that my samples were prepared properly.
I've seen all of these problems occur and worked my way around them. However, I could only work my way around the problem because I had enough context to knew what was going wrong. The deeper you head down the raw data chain, the more context you lose and the easier it becomes to make the wrong assumptions. I know that I have one data set that provides pretty damn clear evidence that we violated the conservation of energy. Obviously we didn't, but looking at the data won't tell you that unless you have information on the capacitance of the electrical interconnects in our power supplies on that particular day.
Research should be verifiable and reproducible. However, an order of magnitude in verifiability isn't as useful as an incremental increase in reproducibility. I'd be happy to let every person on earth examine every layer of my data procedure to see if I've made any mistakes, but even I won't fully trust my results until someone repeats the experiment.
The past and current community standard is that is not absolutely necessary to share data, so people run projects for which it is prohibitively expensive to copy the data.
If we chose to weight sharing more highly, we could budget for sharing and/or scale project data to make it more sharable. We might decide that if the data are too expensive to share in practice, we don't fund the project. We choose. It's not impossible.
The best we could come up with was to publish as many details about the simulations as possible and let people run a limited number of simulations on our own servers (see www.gleamviz.org).
Edit: Even so you can't actually redo the original simulations as there have been several updates since, including a complete rewrite of the simulator, updated input databases, etc...
Has anyone seen a deal like this before?
The current system has its faults but some fields are at least making progress ensuring at least some of the data remains available. For certain projects there is currently no feasible way to openly distribute terabytes of data.
Plus how we can be sure there is no some basic mistake in interpretation? And how study can even pass peer review if most of its sources are hidden?
Technical difficulty is not really an excuse. Many studies are based on a few megabytes or even kilobytes of data. Astronomy and physics has no problem to distribute terabytes of data.
To give an example: Tycho de Brahe made lot of measures of planet Mars positions. However he was not very skilled mathematician, so his interpretation would only make orbital parameters more precise. Luckily Kepler who was briliant mathematician had access to his data and could derive Kepler's laws of planetary motion.
Incorrect widely accepted study can cause lot of damage.
I think share-everything is the only long term sustainable way.
I'm all for open data, but "All X must be Y" is often a flawed argument.