People not liking open source (and it's not Oracle)
weblogs.java.net
weblogs.java.net
i think the real reason for not releasing this code is time---you might be surprised at just how much effort goes into making a paper/talk/article polished. sure, you could just release the hacky code you used to do your calculations, but you'd like so much more---for it to be useful, for it to be a credit to you, for it to also do XYZ. research code is efficient in that it gets exactly what you want for the paper; it's rarely usable by anyone else. there is essentially zero payoff in academia in publishing code, and so there's no time to fulfill these desires.
Sharing on something like sourceforge just simple isn't done. It's not because scientists are "cheating," it's because it's not part of the culture.
I am all for opening tools more, but I don't think reproducibility is the reason for it. In any case, reproducibility does not mean everyone using the same software (with the same bugs) - it means independent experiments.
I suspect the extra work involved in publishing code is because the code is in such bad shape that it would decrease confidence in the work that depends on it. If that's the case, then it's all the more important that researchers establish it as a standard, non-optional part of publishing research.
And here's the key incentive: lack of source code can be used as an excuse for ignoring and devaluing competitors' contributions. Scientists may be reluctant to raise a standard that would result in more work for everybody, but they can't resist a means of denigrating their rivals' work. Once the cat is out of the bag, scientists with weak reputations and disruptive results will be forced to publish source code to prevent their work being round-filed. Established scientists will find it untenable to hold themselves to an explicitly lower standard than everybody else, so they will eventually be forced to start publishing their code, too.
You want me to actually change the culture right now in a comment thread on Hacker News? I think predicting how it will happen is the best I can do, sorry.
If that's not clear, you did not explain why academics would "use failure to release source code to throw doubt on the results of their rivals" when they do not do that now. I don't disagree with you that this would be preferable to the status-quo, nor am I disagreeing that it would be sustainable once the change occurred. I just don't see any reason why the change will occur in the first place.
That said, it is often useful to release code for other reasons. For applied areas, releasing the code in the form of a software tool is fairly common. This allows people from the target area, who may not be computer scientists, easy access to the new method.
Code review seems to be nearly non-existent in the scientific community, although it does happen occasionally when one researcher uses another's code base for further experimentation.
Science is discussed and advanced through papers. These papers aren't an exact recording or rendition of what was accomplished, but a discussion of the idea, methods used, and eventually a conclusion. Methods may be algorithms, but it doesn't necessarily mean the actual source code.
The idea is that the actual code that a scientist used isn't terribly valuable. If another scientist wants to verify the conclusions, they should just whip up another bit of code (or, do the same experiment) themselves.
Asking for scientists to contribute source code is not wrong, it just isn't done (to my knowledge) with other modes of scientific exploration.
Why don't we publish our code? Do you publish one-off shell scripts or database migration code? The only important code is already open source or is (unfortunately) commercial in nature.
Maybe there is a barrier to adoption in the form of git. A more polished version of Gist - marketed to scientists and researchers - would increase uptake.
The point is that scientists don't need to publish certain things, one of them being the source code that helped them discover their findings.
Shell scripts and database migrations aren't super interesting, but you would want them if you needed to replicate results. I'm not saying they're interesting as published research, but neither is that random piece of scratch paper you used while trying to understand what is going on in some biological process.
It seems like the problem is that non-scientists don't understand the way science happens, and that often the source code isn't something that you would include in an academic journal.
Should science be done differently? Maybe. But that discussion is different than "why wont they publish their source?"
That's really not an insightful conclusion. I think it's rather a question of pride and effort: the code probably needs to be cleaned up before release, and the reward for investing that time is very low (career-wise). Researchers are going to publish their source code when it becomes required by most major journals and conferences.
Amongst dedicated biologists, there's no motivation to release your source code. It's the job of a biologist to publish papers so that he or she can pull down the next grant. Anything that doesn't further this goal is going to get zero traction.
Publishing it doesn't help you get into Nature or Science. Sure, Sean Eddy (www.psc.edu/general/software/packages/hmmer/) is going to release his code, but he's a methods guy. The software is the focus of publication. The only way code is going to get published is if the journals refuse to publish results without code.