What's the point of writing good scientific software?
bioinformaticszen.com
bioinformaticszen.com
I'm starting to think, there may be a field of study here. I regularly take software that "proves" some hypothesis or "shows some good results", tear it down, and throw some engineering at it, only to find that the results are not reproduced, or the benefits are severely diminished. Then I have to go track down the why... because until I can show it, it is assumed my fault.[1] The results of some of this are probably paper worthy themselves.
I think a useful field, or useful conference at least could be built for people in this typeof postions, studying the meta effects of software on research. How can we report the issues found, or the updates to the numbers, without putting black marks on the reputations of people who actually are doing good work?
Another interesting phenomenon that is worth study is that the refactoring/rewriting process often gets real results, but it turns out the mechanism for the improvement isn't what the original researcher thought/claimed. It is something perhaps related, perhaps a side effect, and so on. There needs to be a way to recognize both the original researcher, the programmer who found the issues, and the follow-up researchers who did some more determination of the problem.
[1] This isn't as antagonistic as it sounds. It is actually a nice check on my own mistakes. Did the differences in what the researcher did and what I did introduce some strange side effect? Did I remove a shortcut that wasn't actually a shortcut and I misunderstood? A hundred other things on both sides... Research has a large component of "we don't know what we're doing, axiomatically so", and as such it is a decent way of finding out more info.
The Effects of FreeSurfer Version, Workstation Type, and Macintosh Operating System Version on Anatomical Volume and Cortical Thickness Measurements
I recently introduced my advisor to Github, and he thought it was a good idea; however, there were a few hesitations. The first, most importantly, is the likeliness of a bug. If you put your code on a very public website like Github, there is a chance it's going to be scrutinized by everyone in your field.
Now, unless you are one of the best programmers who has ever lived, there are bound to be bugs in your software, and when someone discovers them, it could have a deleterious effect on any journal articles you've written that used that code. The issue is that even though most bugs do not lead to significant changes in results, you would still need to redo all of your data to make sure that is the case. The software industry has long recognized buggy software as a reality, but I don't think the scientific community is as tolerant of it (hence the reason a lot of people hide their code).
For my MD simulations, I use the well-known LAMMPS package. Bugs in it are discovered all the time! (http://lammps.sandia.gov/bug.html). So I think there needs to be a collective realization among the scientific community that these are bound to occur and authors of journal articles can't be persecuted all the time for it. A lot of computational work is the art of approximation so I would just lump "human incompetency" under one of those approximation factors.
Despite this risk, I think I'm still going to release my code at some point as I would personally welcome critique and improvement suggestions. I'd like to think I'm a better coder than most scientists since I've been coding since I was twelve in multiple language paradigms and have won a major hackathon, but eh, who knows. I'm quite sure my environment isn't up to industry standards because I've always coded solo rather than in a team.
You would rather leave potentially incorrect work standing than have a bug corrected?
What the f*ck has happened to science.
This is fairly standard in the science industry (I use the term deliberately). I get rated (in nuclear and particle physics) on how many papers I've published, what indictions of recognition I get from my peers and at which conferences I've been invited to speak. Short of something like a Millennium Prize or a Nobel (one-in-a-million, and both still very political), there are almost no direct rewards to accuracy or importance.
Academics tend not to be the people with the most understanding or the best direct insight - they're the people with the most friends on committees and who do the most fashionable things (usually badly). And so they appoint new academics who are a little bit worse than them (no point in opening yourself to competition - instead bring in people who are going to be grateful to you!) and the cycle continues.
The most sickening part of it all is that who are the ones who perpetuate this system? It is us (other academics).
What can we do to change it? I love research. Having almost 10 years of experience but still earning less than if I had gone straight into industry after a masters degree wears you down though.
I think ultimately it is because scientists add little of value to society in the short term. An MD who has a comparable number of years of education can expect to earn 10X that of an equivalent scientist because they value immediately.
I like Feynman's quote "It's a kind of scientific integrity, a principle of scientific thought that corresponds to a kind of utter honesty--a kind of leaning over backwards. For example, if you're doing an experiment, you should report everything that you think might make it invalid--not only what you think is right about it: other causes that could possibly explain your results; and things you thought of that you've eliminated by some other experiment, and how they worked--to make sure the other fellow can tell they have been eliminated." [1]
> It sounds like you don't want to release your code because someone else might find a bug that you missed
I already said at the end of my post that I decided to release my code. Everyone will be able to see it.
> You would rather leave potentially incorrect work standing than have a bug corrected?
Of course not. You missed the whole point of my post. My point was "Here's a situation in science that needs fixing". "People are hesitant to fix it because...". "I'm going to personally work towards a solution."
Let me ask: what incentive does any scientist have at all to publish their code? You're not going to make money off of it. It's in an obscure niche so you're not going to be world-famous with it. You may get citations to your work, but you're just making yourself vulnerable to having your reputation destroyed because of a bug that nullified all of your articles' results. This is why nobody wants to do it. I'm not saying it's right; I'm saying that it's the status quo.
Most big scientific packages are funded by the DOE, NSF, and others. That's likely the only reason they are even out there.
holy moly x 2
How about that it represents a fuller account of what you did and how you did it? (bugs or no bugs). Isn't scientific publishing supposed to be about reporting what you did as accurately as possible so that others can (1) understand and (2) replicate?
BTW the * in my f*ck from above stands for "ra", what did you think it stood for?
> How about that it represents a fuller account of what you did and how you did it?
So what? Journals never require your code to be submitted. It's not going to increase your article's chance of acceptance. And nobody asks for your code anyway. Why should I publish it if it's not going to bring me any benefits?
> Isn't scientific publishing supposed to be about reporting what you did as accurately as possible so that others can (1) understand and (2) replicate?
In an idealized world, yes. But nobody else does it so why should I?
From the perspective of what "science" is claimed to mean, namely the advancement of human knowledge in a way which is repeatable and verifiable, it seems axiomatic that sharing your algorithms, code, and data are necessary and beneficial to the scientific community.
As a researcher, you don't gain a lot from publishing YOUR code ... but you sure might gain a lot from being able to re-use someone else's code in your domain, or more easily replicate someone else's experiment.
In short: You should share your code because it's the Right Thing to Do if you want to grow human knowledge.
It seems that way, but it isn't. If a cultural foolishness causes you to lose significant credibility undeservedly, it's actually more beneficial to withhold things that could damage it so that you can continue your essentially still-very-useful work.
You could spearhead the fight against said cultural foolishness, but that's time and energy spent doing something other than the work you want to do and are best at.
Someone has to, but why you? This is the Bystander Effect.
So, I guess I can see that argument here. Just because the software has a "bug" isn't sufficient to conclude the results are not accurate. The prejudicial value outweighs the probative value.
wow. has it come to this? really? maybe you should reconsider your career choice
Science shouldn't be like this but I think it's zero sum.
You might want to automate the process of re-running your analysis, though. It's a good idea in general, and especially if you anticipate needing to make minor tweaks to the software involved.
In contrast imagine an academic scenario where I'm testing a hypothesis using someone else's software. How do I notice a bug? I'll notice the bugs where the output is in the wrong format for example. More subtle bugs I won't notice because I don't have any expectation on the results. This will affect the conclusion I draw though.
If you find yourself re-using that bit of code, then it may be worth cleaning it up and making it maintainable. If people start sending you requests for it, then it may be worthwhile open sourcing it, documenting it, maintaining it, etc.--but only if you have time.
I do make open source scientific software as part of my job, but I'm at a later stage in my career and it's not something I would have a science postdoc work on--it's just not fair to them and their career prospects within science...
Recently, someone asked for some reduction code that I've developed and I realized that while it was documented, I didn't have time to refactor it and clean it up--finally, I just put it on github and told them to contact me if they had questions--they were happy to have it as a starting point for what they wanted to work on. So, if you believe that you've made something worthwhile, but don't have the bandwidth to maintain it and other people might find it useful, sometimes it might be better to just put it out there and let people play with it--no guarantees, but it may help someone else get started...
You can get a large number of citations in some subfields for writing commonly used software--but it may or may not help your career. For example, I have friends at various institutions around the world that tell me that their management gives them no credit for developing useful software (complete with lectures, updates, documentation, etc.)--they just release it because they feel they should and most of them are also already tenured in their positions.
Good luck!!!
Admittedly, it's tricky to do, since that name recognition only really matters if you can also manage to publish enough papers. Realistically grad school is more likely for that than during a postdoc or as an assistant professor. Some grad students manage to release some widely used software (well, usually "widely" in a particular niche), which I think does help them build up more prominence than someone in their career stage might otherwise have had.
On a different angle, having produced some reasonably decent software can be a nice thing to have in your back pocket if you ever consider moving to industry. Having N papers and one decent software package is probably a better academia-to-industry transition CV than N+2 papers and no software packages.
I also strongly agree with your other point about just making it open source and if anyone needs it they can get download it and ask you questions.
I have however started to realise that perhaps something I thought would be very useful may be of little interest to anyone else. Furthermore the effort I have put into testing and documentation may not have been the best use of my time if no one but I will use it. As my time as a post doc is limited, the extra time effort spent on improving these tools could have instead have been spent elsewhere."
From a purely selfish perspective, I've found that documenting and cleaning up my own code benefits me in the future. Even if it's a one-off, single-purpose utility that I'll never use again in the future, I often find myself needing to borrow bits of code from my old projects. ("Oh, I solved this problem before. How did I do it? Let's dig up that old, old project...") At which point, present-day me benefits if my past self bothered to actually document things and make sure they're reasonably robust.
There are countless other reasons (moral and pragmatic) to document, test, and open-source one's code, of course! Many of them more important than the ability to crib one's old code, I'd argue.
But the author seems to have considered (and discarded) them...
I did this because I wished that all bioinformatics software had this attention to usability and documentation. However now I wonder what was the point of all of this if no one ever ends up using it? I could have done the minimum for publication then spent this time working on finishing other manuscripts I have waiting.
As I wrote though, I agree with what you wrote in your comment. I just don't think there is any incentive for post-docs in academia to prioritise writing good software over pushing out additional papers.
[1]: https://github.com/michaelbarton/genomer
[2]: http://next.gs
[3]: https://github.com/michaelbarton/genomer-plugin-view/tree/ma...
[4]: https://github.com/michaelbarton/chromosome-pfluorescens-r12...
I ought to have written, "But the author has considered them and concluded that their benefits don't outweigh their costs in his case."
I totally support what you're saying! IMHO it's definitely not worth it to document/open-source/etc code at the cost of one's career or happiness, especially when the code is of questionable utility to others.
You aren't building a system for other users, you aren't really doing anything other than one-off analysis to create charts, which will be explained in a paper.
Things have changed somewhat since the early 2000's, but the concept remains the same. Nowadays, for interesting or controversial results other scientists want to be able to verify your results. However, that is usually more related to your data and how you processed it, rather than your software algorithms (which should be explained in the paper, and can be recreated from that).
So do these systems need to have reams of documentation? Probably not. However, if you leave the system for two years and come back to work on it, or figure out how it used to work, then you best have enough commenting with a thorough readme about some of the decisions you made and why. It's more analogous to scripting rather than software engineering.
(I ask this because, for me, I think the former is much more likely than the latter.)
However, there are a ton of bioinformatics libraries released every year, and almost none of them gain any traction. Nature publications are far more frequent than important new libraries, and you need more political clout to get the library popular than you need to get a Nature paper.
Really, you need to gauge usefulness and interest of your library before you devote a ton of time to it. It's a lot like a startup's product.
Is it more important to get the result, or to use other people's code? Reinventing the wheel is a minor sin compared to not getting results.
The key is to find a modular component that is difficult or novel, and at the same time broadly applicable and hopefully extensible.
That. Like someone pointed out, I find that documenting and testing the key parts (that is, those I know at least I will reuse) is always a good investment of my time and prevents major headaches down the road. I've been experimenting with project structures that clearly separates the set of tools and functions that will be reusable, and those that are one shot. I focus all my testing efforts on former, and cut myself some slack on the latter.
Btw, I speak from a "scientist" perspective, and nothing I say applies to professional software engineering (I mean, I don't think it does).
My bona fides: I helped create a widely-used software system in my field, and have received reasonable credit for it as a scientific contribution.
This reminds me: A man was seen cutting a tree down with a dull bladed axe. A bystander asked him "Why not sharpen your axe first?". The cutter responded "I don't have the time!".
Look at RapidMiner (developed at U. Dortmund), Stanford's CoreNLP, and the brat rapid annotation tool. These are better than a lot of commercial tools. They are more text-analytics than bioinformatics, but same diff.
And now I hear this questioning the value of writing up and polishing scientific software!