The unsung heroes of scientific software
nature.com
nature.com
Failure to distribute the data and code not only greatly devalues the contributions of many scientists, it makes replication far more difficult, and opens the door for outright fabrication.
If computation is a necessary element of the research, it ought to be a necessary element of publication as well.
This is also very true for statistical analysis, many papers do not provide R scripts for their analysis and only a brief overview of the analysis.
Reasons why labs don't produce code:
(1) Fear of being scooped; suppose you put in tons of man-hours into a custom population genetics association study for malaria in Africa, you don't want your competitor lab to sequence a bunch samples in SE Asia and run your code as-is and publish a paper when you have the sequencers to do the same.
(2) Fear of competitors not being able to replicate stochastic ML results; some machine-learning and artificial neural network papers applied to Biology are stochastic. Past authors have been accused of "cherry-picking" the most "optimistic" runs that show positive results, e.g., (https://liorpachter.wordpress.com/2014/02/11/the-network-non...)
(3) Focus on science, not tool production; most labs' focus is to produce publications, not open-source software. User base for scientific software is very niche, unlike Web MVC frameworks; making the payoff calculus of packaging and supporting external users not so great. Furthermore, most labs have custom databases making integrating external software, externalizing internal software difficult.
1. "It'll be released soon" on a years old page 2. 404'ing personal pages on university sites 3. Finally some code
But then the code didn't, and couldn't, generate the images they have. So now I'm not actually sure if the algorithm works as they say.
This was also one of the more successful results I've had in trying to find research code outside of ML. I'm sure other disciplines are also good but ML is mostly what I've had to find.
There's a part of the experiment which virtually anyone could replicate for a material cost of near 0 and check in detail if they wanted to and yet that seems to be the most poorly explained and shared part of it. You wouldn't get away with just mentioning the protocol you used in the lab without any references or explanations, would you?
We just wouldn't accept these kind of excuses for other descriptions of methodology, even though they might apply. There's no reason that code should be an exception.
> We didn't release our first, second or third papers because we didn't want to be scooped on the fourth and if we publish then someone might build on our work!
Which happens all the time. If you are working on something big you'll sometimes hold back on publishing parts of your research until you've completed the big thing for just that reason. Then once you've completed the big thing you publish all four papers either together or in quick succession.
That sounds like an ML phrasing for "fear of non-reproducibility", which is basically a fear of their research being bad.
This prevents a lot of good research from coming out. We should not just rely on publish or perish. Ifa lab makes a net contribution, they should still be rewarded with grants even if the idea gets "scooped".
(1) use an off the shelf prng
(2) provide the data
(3) provide the seed
http://software-carpentry.org/
I took their instructor training last summer but I haven't had a chance to run a workshop yet. I think it's a great idea, especially given my recent exposure to researcher-written codebases.
(Using eclipse seems problematic when there is a text editor that doesn't involve any cognitive overhead to many bioinformaticians. The step through debugger did seem to appeal to this one though).
That is, it seems that most of the argument for "open code" in the scientific review sense implies that 1) the software should be available at no cost, and 2) there is no need to support commercialization of said software. These principles are different than the four freedoms of "free software", and I have difficulties in reconciling the differences.
For a concrete example, I am self-employed. I sell scientific software. All of my customers receive it under the BSD license, after they pay me a good chunk of money. Thus, I sell free software for scientific research.
The FSF says "Selling a copy of a free program is legitimate, and we encourage it. (Quoting https://www.gnu.org/philosophy/selling.html .) But most of the time when I see people say "open code" or something similar, they want the right to view, use, test, and modify the source code at no cost.
If I publish a peer-reviewed paper about the software, should I be required to distribute the software to readers for no cost? Or may I set a fee of, say, $25,000 to get access to the source code under a free license? (I'm well aware of the loophole where I could publish that something is "open", but require a payment of $1 billion. My question is, what is a reasonable and fair price to charge?)
On the flip side, if I am required to publish my software under a free/open source license and not charge for access to the code, then that means any reader can take my source code and commercialize it, or even simply give it away. Commercialization is part of the four freedoms of free software, but it ends up reducing my market size and knocking other parts of the four freedoms. I won't have as much money to continue my self-funded development and research.
The principles are important because they help set guidelines for other questions. Can I require that people register their use of the software before getting access to a no-cost copy? Can I wait 6 months to respond to those requests? How long am I required to host the software, or will the journal manage all of that? May I include non-free license terms, like a requirement to cite X if someone uses the software? Does minimized/obsfucated code suffice? And many more details that have been resolved in the context of F/OSS software but have not, I believe, been resolve in the context of what's needed for peer reviewed publications.
There are two issues I see with closed source research software.
It seems to me like someone outside the research group should review all code run for a paper as part of peer review. If there are glaring off-by-one errors, race conditions, etc., then why should we trust the results produced by the software? There are a lot of things that should get caught in peer review that would only be red flags if someone experienced was able to look at the program source.
Also, I've always been under the impression that science should be reproducible. If no one can replicate your results, then how can we trust your findings? While this is technically possible without access to source code, closing source code off from the community really makes reproducing results harder.
I'm not sure that it would be fair for journals to require open source availability for publishing, but if they don't then we need a creative solution because these are real problems.
Consider the X-PLOR program for crystallography refinement, where an academic/ research license was a few hundred dollars. That came with source code, plus the right to distribute patches and other modifications to anyone else who had an X-PLOR license, but not to everyone else.
Consider the NAUTY program for graph isomorphism, from http://users.cecs.anu.edu.au/~bdm/nauty/ . The license says "Permission is hereby given for use and/or distribution with the exception of sale for profit or application with nontrivial military significance."
Consider the many programs available for free and unrestricted download from university web sites which are "for academic use only", some of which require users to cite a given paper. (Eg, http://www.maths.lth.se/matematiklth/personal/sminchis/code/... is one I easily found with a web search which is available in source code and is "free of charge for non-commercial research and education purposes".)
These are neither open source nor free, so are they "closed source research software"? If "closed source" means "not open source" - which is the usual view - then yes, the above programs are all closed source.
Yet all of them are available for peer review, reproducibility studies, etc. that you want.
While on the other hand, to get what you want requires principles different than what the Free Software Foundation considers to be one of the essential freedoms in programming - the freedom to sell software. So at the very least, "peer-verifiable" software, for lack of a better term, is not compatible with "free software."
I personally think it's unwise to even use the terms "open source" and "closed source" in this discussion because of the confusion it adds. However, as most researchers come from academic or government labs with non-commercial funding sources, I can see why self-funded, for-profit scientific software may be overlooked.
Perhaps I'm being too loose in my terminology, but in this context I used "closed source" to mean "source code unavailable to the community." In my view the attached license(s) matter a lot less than making it so that the science can be reviewed, reproduced, and improved upon.
If the source code can be reviewed for free but not redistributed, reused or commercialized for free, then I don't see why that would hinder endeavors to review and reproduce the research. However having to go through byzantine and expensive processes to procure source code would be an impediment to those in academia who don't have funding to pay for licenses to source code just to review a paper, for example. Maybe I'm not reading closely enough, but that sounds a lot like what you're advocating in your original comment.
What I don't understand is how to set up practical guidelines.
For example, consider "reviewed for free but not redistributed". If I review the software, and find an error, what do I do? Should I publish a paper which demonstrates the difference between the original and corrected versions? If so, I need to include the fixed code, and perhaps also the original. But that's a redistribution.
"would be an impediment to those in academia who don't have funding to pay for licenses to source code"
As a minor point which is big in my mind - most academic groups have more funding to pay for licenses than I, a self-funded, for-profit researcher, have.
"but that sounds a lot like what you're advocating in your original comment"
I mentioned that, to explain the view of most people who want access the source code. I was not advocating it.
My question was, is this requirement important enough that all of the source code must be made available at no cost? If so, it's in opposition to the FSF's four freedoms, which encourages people to sell free software, so there must be some other philosophical underpinning to justify the no-cost argument.
What is that philosophy? It can't simply be "to verify" because there are some problems, like factoring RSA-360:
2186820202343172631466406372285792654649158564828384065217121866374227745448
7764963889680817334211643637752157994969516984539482486678141304751672197524
0052350576247238785129338002757406892629970748212734663781952170745916609168
9358372359962787832802257421757011302526265184263565623426823456522539874717
61591019113926725623095606566457918240614767013806590649
where it's trivial to verify the solution is correct without reproducing the calculations.And what counts as a "byzantine and expensive processes"? The process of reproducing one of the CERN papers, especially if I need to make my own accelerator, is non-trivial and expensive. Some molecular dynamics simulation software only runs on custom-made ASIC hardware, or uses $100K+ of CPU time. A software cost of $20K is only a small part of the overall cost in that case.
Who pays the developers? If you aren't paying the developers much, what exactly are they going to catch in the 2-4 hours they may have to look at it?
I wish I had a good solution; I think about the best we can do is community-pooled development efforts ala openfoam, bioconductor, numpy, sklearn.
Where it is easy to make it easy to replicate, it should be made easy. Releasing source code, even in unsanitized form, allows for much easier inspection and replication.
The research world is also much different today than a century ago. Today the field is crowded with researchers who must publish or perish. The speed of papers, many of which are in fact nonsense, is churning out is unprecedented. As a result, a paper that can be verified but is hard to verified is often never properly verified at all.
No, but making it easy as reasonably possible or at least not going out of your way to make it hard would be nice.
What I dislike the most are computational articles that give no indication whatsoever about the employed tools and programming languages.
If I develop my own tools, which no one else has, and I never distribute them, then it's the trivial edge case that everyone who owns those tools (me!) can replicate the results.
If I sell the tools under a BSD license for $1,000,000 then it might still be trivial for those who pay me that sum to reproduce the results. But those who argue for source code access usually want the source code available for $0 or a pittance compared to the development costs.
You agree that it should not be "prohibitively high". How do we turn that into something actionable? If $1M is too high, then what about $100K?
Does the requirement extend to providing documentation? Even if such documentation doesn't already exist? For one job, I fixed a few bugs in software that was commented in Russian, and I speak no Russian. Is this too high of a barrier to entry? And if so, should all code be commented in English?
Correct. Imagine reading an article in a journal that said, "We were able to prove that Theorem 3.1 is true. The proof is omitted because we want to use similar techniques to prove theorems in the future, and by not publishing any details of our proof, we will have an advantage over rival researchers. That will allow us to recover the significant investment of time that we put into doing this proof."
"Using the argument of [1] and [2] it can be shown, with brief step x, it can be shown that this theorem holds in xxxx" and then moving onto the conclusions, is exactly the same as "Using the software developed in [3] to solve Eq. 1, we show that our result is statistically significant."
If the claim is probable, it will and does get through review.
What should be done in the case of when a researcher creates a private company as a result of publicly-funded research, but the research isn't fully released or is obfuscated?
Personally though I think we should give up on papers and everything should be published on an ipython notebook style page where others can play with data and code.
I think that if we started thinking about research software as a research contribution in itself, it would be a good way to accomplish what you talk about, e.g. by making software (and data) something that is published and cited, rather than attached as a supplement to every publication.
In short: benefits for scientists (career, prestige) are misaligned with benefits for science and society (reproducibility, progress, openness).
http://www.artifact-eval.org/ https://plus.google.com/+JanVitek/posts/19BH96G8rrw
The artifact evaluation workflow is a reasonable starting point for what you suggest.
As one example, over a decade ago I wrote a biological simulation program called CompuCell3D it was the successor to CompuCell, a 2D simulator of cells as cellular automata. Since then many papers have been written based on research utilising this software. Not only have I not been credited in any of these publications but my name has been removed from the software, the website never mentions my original contribution and the current maintainers of said software are not responding to my emails.
Granted this code has changed a lot over the years but there are still large parts of the code which are still verbatim from my original software. The researchers in this project see it as totally irrelevant that I wrote the original code because they value code creation MUCH less than their research.
This particular case is especially egregious but it is a good anecdotal example of the problem. Sometimes code is not as significant or as breakthrough as raw research. It really matters what kind of coding you are talking about. You don't necessarily credit the construction crew when dedicating a new building but you do credit the architect. Coders are under valued in today's research environment.
Exactly because it is a research environment. Coders are valued in a tech company environment because that's where they're the stars. In any other organization, such as a transportation company or a government bureau, a coder is just an assistant to the main tasks, and there is no reason it shouldn't be different.
I've noticed this being a good way to maximize credit.
Although both BSD and GPL already require attribution, so you could say that it's redundant.
When I was a Research Assistant/Experimental Officer at a world leading Rnd organization I was paid about 1/3rd of what other jobs with similar entry requirements did.
I think a decent software development capability is one among the many skills a scientist must be able to be comfortable with in order to provide society with valuable research results.
Professional developers are given a task, or a goal, and are focused on the right way to do it. Some do wonders, but they aren't paid to go beyond the goal they are assigned.
When you're a scientist, the software is not a goal, but only a tool that will be subjected to further iterative refinement.
No one else but you knows about the proper efficiency / flexibility / level of abstraction you need. There's a strong need for scientists combining both scientific and development skills.
Btw, I'm looking for a new position ^^
The website isn't required to list your contribution. That would be a courtesy.
Also, if you have a problem with this where you think it's important to change the website or get another form of credit, your lawyer should be talking to your ex-employers lawyer. That should have been clear when the current maintainers failed to respond to your email.
I can't see that article online without paying, but I am certain it cites CompuCell, a Multi-Model Framework For Simulation of Morphogenesis – J. A. Izaguirre, R. Chaturvedi, C. Huang, T. Cickovski, J. Coffland, G. Thomas, G. Forgacs, M. Alber, G. Hentschel, S. A. Newman, and J. A. Glazier, Bioinformatics 20: 1129-1137 (2004)
that original article includes your name on the author list.
So your statement "I have not been credited in any of these publications" isn't really correct; your contribution is credited indirectly through citations. I don't see any problem with this; I doubt papers citing CompuCell3D have any text like "We acknowledge the efforts of <so and so>".
But what happens when the more senior authors (who tend to be experimentalists) get the manuscript, they want to add a bunch of experimental citations and they see citations of computational methods as irrelevant and able to be sacrificed if there are too many citations.
And as you say, the growing popularity of GitHub gives us all kinds of cool data even when there's no central package manager for the language. In fact, we're mining imports of every Python and R project on GitHub right now to build out the dependency network beyond the (much much smaller) CRAN and PyPi networks.
The idea with Depsy has been to launch quickly with two languages, so people could see what it looks like, then iterate and add more as we get feedback. So we'll count your comment as +1 for C and C++ :)
You have to start somewhere, and if you get GitHub working then it would be (hopefully) much easier to then include BitBucket, GitLab, etc. Indexing self-hosted repos would be pretty tricky, I imagine, if only because you'd then need to maintain a list of all of the servers to clone from.
Further, there's also the challenge of including software from all of the research groups who don't even appear to use version control, or if they do it's locked away on a private server or service. How does one attribute authorship rights to those people without source history?
Anyways, this is a long ramble now, but I'm mostly trying to illustrate that it's difficult to do what Depsy does with software when it's written in languages without a canonical package repository. That doesn't mean they shouldn't try to expand their reach just because "GitHub isn't enough."
Of course, if it's not a public server then you can't do anything, but then again in that case the authors can't really complain they don't get any credit for what they do.
[1] www.hypnagogic.net/rob/ [2] www.knotplot.com
1) it's actually harder to get reproducible computational experiments than they expected. For example, you can run same VM on a different processor and get different results, which makes bitwise reproduction hard, and statistical tests for nonbitwise equality are harder
2) developing and maintaining the VMs and the environment takes a fair amount of effort from skilled people
3) the resulting improvements to science don't seem to exceed the cost thresholds implied by #1 and #2, and nobody's volunteering their time.
My conclusion: good idea, but probably not critically necessary.