Publish your computer code: it is good enough
nature.com
nature.com
I've seen source code for software used to produce results in major publications like Nature that was so poor quality it is surprising it even compiled. You don't generally get tenure for producing well written software, so there needs to be another incentive for scientists to spend the time to write well thought out and well documented code. I think sunlight helps tremendously.
Warren DeLano, the author of PyMol, the very popular molecular visualization software, was an early voice about the importance of open source software in the sciences. Unfortunately, he is no longer with us, but it is hearting to see others now taking up the cause.
The only way to publish software in a scientifically robust manner is to share source code, and that means publishing via the internet in an open-access/open-source fashion. - Warren DeLano (2005)
It would be better to describe your methods, so someone else can implement it in their preferred tools. If they still replicate your results, that strikes me as being much stronger than just rerunning the same, possibly buggy, code.
I mean, if you're working with monkeys, other scientists can't reuse your monkey. They have to follow the procedures you describe on their own animals.
Another issue is that the code might not be very useful to other labs as-is. The code might be for unique, custom-made hardware, or an unusual configuration of equipment, or in an unusual language.
I get the impression that software types are inclined to put way too much significance on the source code. I'm not surprised the author of the linked item in Nature is a software engineer.
I wonder, did scientists in the 1930s publish their scratchpads full of calculations?
Science can be done in isolation. Robinson Crusoe could do science on his small island in the tropics. Without peer review or replication, you're just not doing the particular kind of science that's most beneficial to society: the kind that can be built upon.
A man working alone is doing science, he's just not benefiting society as well as he might. Science isn't "pushing the bounds of human knowledge." Science is forming hypotheses and testing them empirically. When kids drop two objects of different weights to see which falls faster, they're doing science just as much as Galileo was, despite the fact that their findings will never get published.
What a researcher does doesn't magically become "science" when he publishes. It's science all along. If she does it in a cave for years before she publishes, it's still science for all those years she was in her cave.
Children who perform experiments that have been done a million times before are still being scientists.
People who perform experiments but never get published are still scientists.
Science is about the process, it's not about the result.
I am a proponent of open notebook science, and strongly believe that the benefits outweigh the negatives. Besides the feel-good arguments that this advances science and reproducibility, I'll point out a selfish motivation for releasing code: It makes it more likely that other researchers actually try your methodology and cite you.
I used to get mired down in trying to do formal releases. Now, I realized that releases are a hindrance, and by default all my research code lives in a github repo from day one:
For example, here is the page I published on my word representation research, with links to my github code: http://metaoptimize.com/projects/wordreprs/
You should point out that refusing to publish code is like being intentionally hazy about your experimental protocol.
When I started, code was jealously guarded as a "secret weapon" in the global competition for publishable results. Then the MILC group (http://physics.indiana.edu/~sg/milc.html) started releasing both their data and their code. The result was that they were widely cited and gained much respect in the scientific community, becoming a de facto standard, and facilitated research at institutions without the resources necessary for such computational preliminaries.
I'm pleased to say that this approach has become more the norm in this field at least, encouraged by SciDAC ( http://www.scidac.gov/) for example, with raw data (e.g. http://qcd.nersc.gov/, http://www.gridpp.ac.uk/qcdgrid/) and code (e.g. http://usqcd.jlab.org/usqcd-docs/chroma/, http://fermiqcd.net) being routinely made available.
The result is more science all round, which can only be a good thing. The "scooping" thing we all worried about turns out not be such an issue after all.
Even worse are the researchers who keep their models and data sets secret due to paranoia that some colleague will publish first, which admittedly does happen.
Cultivating a spirit of genuine cooperation and sharing in academia would do wonders for the progress of the sciences, but there are so many hurdles that need to be removed. It's not just a matter of knowing that it's good to release source code or feeling confident in including it as the article suggests.
Here's a good page encouraging everyone to boycott them: http://mingus.as.arizona.edu/~bjw/software/boycottnr.html
The biggest concerns I've heard from other researchers (ie, professors) has to do with being beaten at a publish. I think that as long as your paper itself isn't easy to find pre-submission, then you're okay, especially if nobody will really understand what you're doing anyway. So, I'm not too nervous about the prospect. It's just code for now.
Unfortunately, even if scientists release their own code, it's often just a small part of the big picture. In engineering, at least, MATLAB, Mathematica and commercial Finite Element packages are everywhere. My own project uses MATLAB and COMSOL in tandem, meaning that what I wrote only really serves to glue a bunch of completely closed algorithms together. Personally, I'd love to see a completely open, usable and documented FEA stack (which would ideally include an adaptive mesher, some FEM algorithms, a post-process visualizer, and both a gui and a not-shitty scripting API).
Did you have to clear this with your university/advisor? AFAIK, my university has copyright on code I produce on paid time or using university facilities/equipment, which covers my research and even a good chunk of my homework.
meaning that what I wrote only really serves to glue a bunch of completely closed algorithms together
This is consistent with what I've run across. It seems a lot of research revolves around modifications to an existing system, but in CS, it seems the established code base is more likely to be open source (e.g. Jikes RVM) or be proprietary but still have source code readily available (e.g. SimpleScalar).
I don't know if I had to, but I did discuss it with my advisor, mostly due to the "someone could steal my thunder" issue, and we were basically in agreement. I actually have a really cool advisor--I lucked out there.
They're very unlikely to stop you if you just go and do it. It only really becomes an issue if you try to make money off of it.
The advantages still outweigh the disadvantages, but it's important to remember that there are always trade-offs.
and that brings us back to why many scientists don't publish at all now...
As long as not publishing code is standard, there's a incentive to do it since it might open you to career-withering criticism...
I have tried to explain (With limited success) to my colleagues in bioinformatics that not unit-testing your code is like using an instrument that hasn't been calibrated, and blindly trusting the results to be right.
This is the standard procedure for when you've developed a complex algorithm to do something novel, and actually working through it by hand would take hours.
I should know, this is how I used to do it before I knew better, and as a result one of my published papers has (very slight) numerical errors in.
The quality issues aren't all that different than a web framework, either. Release one or a small number of code bases, and have everybody pound on and improve them, and you might get somewhere. Have everybody write their own code bases from scratch every time and you'll get yourself the scientific equivalent of http://osvdb.org/search?search[vuln_title]=php&search[te... .
To be honest at this point when I see a news article that says anything about a "computer model" I almost immediately tune out. The exception is that I read for some sign that the model has been verified against the real world; for instance, protein folding models don't bother me for that reason. But this is the exception, not the rule. When it became acceptable "science" to build immense computer models with no particular need to check them against reality before running off and announcing immense world-shattering results I'm not exactly sure, but it was a great loss to humanity.
A lot of time might be wasted going down that path too.
In the real world, most scientists will keep modifying their program until it gets an expected result (either because the program is correct, or because multiple mistakes cancel each-other out, or the standard theory is wrong but they feel obliged to keep debugging until they can prove it to be correct). Sure, there's times when a computational model proves something interesting (unexpected), but then the researchers may have to open source it just to prove they didn't screw it up.
A lot of code is going to be like that: applying known functions that can be identified by name. If a paper says it used an FFT on some data, there are plenty of well-tested FFT implementations that can be used by another researcher to try to replicate the original result. The original researcher's FFT code, if any, isn't really necessary as long as the function, the inputs, and the outputs are well-documented.
python.genedrift.org
Another issue well as code quality, transparency, reproducibility etc. is simply reuse. There's a lot of wasted effort in academia where people are constantly reimplementing simple things from scratch that do basically the same thing as their peers' code does.
Okay, this is true in other fields too, but in our field it's public money getting wasted.
http://cms.mpi.univie.ac.at/vasp/vasp/How_obtain_VASP_packag...
Due to advances in computer literacy and the creation of competing projects, the landscape has been changing in recent years:
http://en.wikipedia.org/wiki/List_of_software_for_molecular_...
http://en.wikipedia.org/wiki/Quantum_chemistry_computer_prog...
The code actually corresponding to a given paper is likely to be a small amount that builds on packages such as those above, or products like Matlab, whatever. There could be code to implement an experiment, and Matlab code to analyze the data.
I agree that nowadays the fear of harming one's reputation through abuse of their code sounds unreasonable. Leaders of major scientific codes told me, however, that this was their overwriting concern in the 80s and 90s. As an alternative to imposing barriers to entry through fees, some groups demanded collaboration on the first project that used their code (a method that doesn't scale well, but doesn't require writing of thorough documentation, another thing that many scientists dislike). Hopefully, the times they are a-changing.
1. In the absence of a well defined standard, it’s the individual scientist’s/consortium’s responsibility to define and actively use an organized meta-data standard.
2. If it’s not open source it’s not science.
3. A snapshot of the source code used to generate results should be given/pointed to when the results are presented.
4. Minimizing reproduction time is an integral part of science.
5. Principles 1-4 should be be demanded, by funding agencies, program heads, and research advisors.Also, in certain disciplines, not only papers, but also the software used to obtain the published result is peer reviewed. For instance, the journal Mathematical Programming Computation (http://www2.isye.gatech.edu/~wcook/mpc/index.html) accepts papers accompanied by the software used by the authors, which is tested and reviewed by technical editors.
http://www.econometricsociety.org/submissions.asp#Replicatio...