Majority of published scientific data not recoverable 20 years later
upi.com
upi.com
This doesn't solve the problem for research where the datasets are in the many terabytes, but then again there are many papers where the datasets are well under a gigabyte.
No one got tenure for a well curated data set.
We've had a ton of scientists sacrificing their own individual reputation and relationship with publishers (and of their groups) in name of something everyone agrees but few really stand up for, because the personal gains are almost exclusively negative. That's textbook use case of regulation.
I don't know specifically what should be done but perhaps a requiring publishers certain obligations (e.g. responsibility of maintaining papers for a long date and turning them public afterwards); or maybe simply a universal obligation to open publications after say 5 years.
People fear this will compromise quality or sustainability of publishers. But the community need publishers. It's a tag of credibility. So if publishers are into trouble (and they're really needed) they'll find a way by e.g. demanding payments from publications from the most wealthy labs.
The problem is that whether or not a paper is Open Access, "Data is available on request from the author" may be an undocumented bit of spaghetti code, may be stored on a Zip disk around here somewhere I'm sure, or may just be lost.
Regulating "You must make your data accessible, and maintain it well" is much harder to implement, and much harder to check. Some grants now have sections describing what will happen to the data etc., but right now there really is very little reason beyond their own personal desire for researchers to maintain good quality software and data repositories.
Additionally, the discussion over shared data is usually a cry to improve reproducibility. In this publication-centric world, one can't publish a paper that is just "we reproduced an existing paper", so outsiders look to "poach" new applications or findings from the data.
The data from my dissertation is on a Zip disk. I didn't mean for it to get lost to the world, but, you know...
I'm on a project now where provision of the raw data to the granting agency for public availability is a requirement. But the metadata, database structuring, and answering questions from people who take my data and have a question about it has added noticeably to my workload.
I know this is an imposition under current incentive structures but it seems to me on a macro level that this is exactly what we the public would want to see occurring, people actively reading other people's research and data and unrelated scientists asking questions of each other and reviewing each others findings.
However, this was simulated data, it could be recreated by checking out some old version of the code and rerunning it. These days, it wouldn't even take that long. The situation is different where someone's made measurements of the real world, since those are truly irreplaceable.
Study without original raw data or source code, is just authors opinion, not a science!
The current system has its faults but some fields are at least making progress ensuring at least some of the data remains available. For certain projects there is currently no feasible way to openly distribute terabytes of data.
Plus how we can be sure there is no some basic mistake in interpretation? And how study can even pass peer review if most of its sources are hidden?
Technical difficulty is not really an excuse. Many studies are based on a few megabytes or even kilobytes of data. Astronomy and physics has no problem to distribute terabytes of data.
To give an example: Tycho de Brahe made lot of measures of planet Mars positions. However he was not very skilled mathematician, so his interpretation would only make orbital parameters more precise. Luckily Kepler who was briliant mathematician had access to his data and could derive Kepler's laws of planetary motion.
Incorrect widely accepted study can cause lot of damage.
I think share-everything is the only long term sustainable way.
The past and current community standard is that is not absolutely necessary to share data, so people run projects for which it is prohibitively expensive to copy the data.
If we chose to weight sharing more highly, we could budget for sharing and/or scale project data to make it more sharable. We might decide that if the data are too expensive to share in practice, we don't fund the project. We choose. It's not impossible.
The best we could come up with was to publish as many details about the simulations as possible and let people run a limited number of simulations on our own servers (see www.gleamviz.org).
Edit: Even so you can't actually redo the original simulations as there have been several updates since, including a complete rewrite of the simulator, updated input databases, etc...
Has anyone seen a deal like this before?
For my thesis, I measured the Paterson function for a series of colloids. I can imagine other scientists finding this useful and I'd be happy to submit it. However, it's not the raw data. What I actually measured is the polarization of a neutron beam, which I then mathematically converted into the Patterson function. So I should probably submit the neutron polarization I measured, so that other scientists can check my transformation. Except that I can't directly measure the polarization - all I really measure are neutron counts versus wavelength for two different spin states, so that must be my raw data. But those counts versus wavelengths are really a histogram of time coded neutron events. And those time coded neutron events are really just voltage spikes out of a signal amplifier and a high speed clock.
If a colleague sent me her voltage spikes, I'd I'd assume she was an idiot and never talk to her again. Yet, I've also see experiments fail because of problems on each of these abstraction layers. The discriminator windows were set improperly, so the voltage spikes didn't correspond to real neutron events. The detector's position had changed, so the time coded neutron events didn't correspond to the neutron wavelengths in the histogram. A magnetic field was pointed in the wrong direction, so the neutron histograms didn't give the real polarization. There was a flaw in the polarization analyzer, so the neutron polarization didn't give the true Patterson function. And all of this is assuming that my samples were prepared properly.
I've seen all of these problems occur and worked my way around them. However, I could only work my way around the problem because I had enough context to knew what was going wrong. The deeper you head down the raw data chain, the more context you lose and the easier it becomes to make the wrong assumptions. I know that I have one data set that provides pretty damn clear evidence that we violated the conservation of energy. Obviously we didn't, but looking at the data won't tell you that unless you have information on the capacitance of the electrical interconnects in our power supplies on that particular day.
Research should be verifiable and reproducible. However, an order of magnitude in verifiability isn't as useful as an incremental increase in reproducibility. I'd be happy to let every person on earth examine every layer of my data procedure to see if I've made any mistakes, but even I won't fully trust my results until someone repeats the experiment.
I'm all for open data, but "All X must be Y" is often a flawed argument.
The diagram is based on one from an earlier paper: 'Nongeospatial Metadata for the Ecological Sciences', 1997, Michener et al. http://dx.doi.org/10.1890/1051-0761(1997)007%5B0330:NMFTES%5...
"With the emergence of online publishing, opportunities to maximize transparency of scientific research have grown considerably. However, these possibilities are still only marginally used. We argue for the implementation of (1) peer-reviewed peer review, (2) transparent editorial hierarchies, and (3) online data publication. First, peer-reviewed peer review entails a community-wide review system in which reviews are published online and rated by peers. This ensures accountability of reviewers, thereby increasing academic quality of reviews. Second, reviewers who write many highly regarded reviews may move to higher editorial positions. Third, online publication of data ensures the possibility of independent verification of inferential claims in published papers. This counters statistical errors and overly positive reporting of statistical results. We illustrate the benefits of these strategies by discussing an example in which the classical publication system has gone awry, namely controversial IQ research. We argue that this case would have likely been avoided using more transparent publication practices. We argue that the proposed system leads to better reviews, meritocratic editorial hierarchies, and a higher degree of replicability of statistical analyses."
Wicherts has published another article, "Publish (Your Data) or (Let the Data) Perish! Why Not Publish Your Data Too?"[2] on how important it is to make data available to other researchers. Wicherts does a lot of research on this issue to try to reduce the number of dubious publications in his main discipline, the psychology of human intelligence. When I see a new publication of primary research in that discipline, I don't take it seriously at all as a description of the facts of the world until I have read that independent researchers have examined the first author's data and found that they check out. Often the data are unavailable, or were misanalyzed in the first place.
[1] Jelte M. Wicherts, Rogier A. Kievit, Marjan Bakker and Denny Borsboom. Letting the daylight in: reviewing the reviewers and other ways to maximize transparency in science. Front. Comput. Neurosci., 03 April 2012 doi: 10.3389/fncom.2012.00020
http://www.frontiersin.org/Computational_Neuroscience/10.338...
[2] Wicherts, J.M. & Bakker, M. (2012). Publish (your data) or (let the data) perish! Why not publish your data too? Intelligence,40, 73-76.
One clear user case is that recently it was published in the New England (a tier one med publication) how all the studies about high blood pressure and salt relationship have their origin in an old animal study with rabbits (If I recall correctly). They fed them with the equivalent for humans of hundreds of grams of salt, as the blood pressure rise it was deduced that salt causes high blood pressure. They also did a meta-study about the relationship of high blood pressure and salt and they didn't find a clear correlation. Maybe this is correct maybe it´s not, but that a medical truth as established as salt=high blood pressure can not be properly traced and known, just let's you see how it's all broken.
edit: typos and spelling