How to support open-source software and stay sane
nature.com
nature.com
But in my field, in most cases the source code is never released at all. That's a far bigger problem that not having support to use it.
So fellow academics, please don't use this article as an excuse to not release your code. When in doubt, just push the thing to Gitlab as is, add a README that says "This is research code for paper X and it is unmaintained.", and disappear. It's not ideal, it's not the best way to do science -- but it's much better than not releasing your code.
Related: the CRAPL http://matt.might.net/articles/crapl/
I agree but I think it's good for people to know that almost any moderately successful open project will attract whiners, complainers, entitled people, and in some cases outright abuse.
From the perspective of society, having no open code is worse than having open code that's unmaintained. We're agreed here.
From an individual contributor's perspective, opening yourself up to varying forms of whining and abuse "for the social good" (and not, say, for your tenure or publication count or whatever a researcher cares about in the moment) is a bigger problem than just sitting quietly on stuff you don't want to become a drain on your life.
It's of course nothing like actual abuse or threats of violence but it's an emotional toll that I certainly wouldn't want to pay. I have better things to do with my life.
Given that precaution, nothing! Sadly a lot of people don't take this precaution.
It's also poor for people distributing code under the license. It jokes about how the program was constructed without thought or design, but that's the very basis of the program's copyright protection. If you ever try to enforce the license, you might regret having to spend time arguing the implications of that joke in court.
Nobody should ever use CRAPL. Use the MIT or GPL licenses and set expectations for maintenance or support to zero in the README as you said.
Previous discussion: https://news.ycombinator.com/item?id=9670497
The sad part is that to a lot of scientists and researchers, software/software engineers isn't something worth paying for. It's not uncommon to see "programmer" jobs that are looking for 3+ years of experience that offer <$15 dollars an hour in the US. Sometimes they're "volunteer intern" positions. Of course the people who end up filling these positions aren't usually actual developers, so the software gets built poorly, eventually gets scrapped, and the cycle continues.
Management also hasn't really evolved past the 90's. Non-technical scientists often want 100% of control and to make each decision, but don't want to spend any time on it. This means developers often have little to no specs to work with, but spend all of their time guessing about what the scientists want, and having to go back and fix everything after.
>“That’s really the tragedy of the funding agencies in general,” says Carpenter. “They’ll fund 50 different groups to make 50 different algorithms, but they won’t pay for one software engineer.”
This is the crux of my frustration. It's not even 50 different algorithms often. A lot of the time, 50 different research groups will be working on very similar programs, and none will be able to deliver a working version.
Though the article mentions that research funding does exist, clicking on one of those funding pages and looking through their examples reveals that only ~1/10 of their websites are actually still active, and they aren't old sites. Again this goes back to the whole "scientists don't value software thing". I've seen scientists happily sign off on spending $20,000+ on hardware components that would usually cost <$100 to make, but balk at contributing $50 yearly to support open source.
I got lucky that I managed to find a place where I get paid fairly, and my boss is actually technical and can manage tech projects well, but these places are few and far between.
Plus, with only a single software engineer, there's a good chance you get unlucky and end up with someone clueless/lazy. You would probably need 3-4 software engineers to make a functioning team with best practices and hedge your bets against accidentally hiring someone who sucks. So now we're talking 10+ grad students.
Open source software is a bit different because many labs can band together to fund things they find useful. But again there are still issues with cost-effectiveness. I'm guessing most lab contributors to OSS would want some sort of quid-pro-quo which may not be realistic for all OSS projects. And by funding OSS you are also funding competing labs' abilities to use the same features you use, which is good for science in general but not good for people's careers sometimes
Note that I said should there. How to write maintainable programs seems to be lacking in research area.
And still, you're getting a generic CS undergrad's caliber of work and responsibility which I would say on average is not great. They might not be as familiar with version control, best practices, etc. and could just end up writing code just as bad as the scientists.
I think if hiring a team you would need at least one somewhat experienced full-time software engineer to act as team lead/PM for the other developers, whether fulltime or students.
For sure, talented undergrad CS developers exist, but in my experience there are far fewer of them than many people might think. Experience counts for a lot, especially when trying to deliver even average quality software work.
As for something made by undergrads that is “maintainable” to a standard broadly comparable with an experienced developer? Maybe if you get lucky...
The fact that they need the experience should already tell you that they're not yet competent at what they are needed for. They should be put on low-impact projects that don't matter, or low-impact projects that do matter with solid mentoring. They should not be made to write software for something that is high-impact for you with little to no mentoring.
The problem is more pay gap with industry. Even though they tend to pay well by academic standards (e.g. no PhD required, yet pay is much higher than a senior postdoc), the salary is still way below what industry in London or Oxbridge offers.
Furthermore, you tend to be surrounded by non-technical people which may be tough in the long term. Nobody appreciates what you do. Not even your boss, who may know zero about computers.
The bottom line is that positions end up being vacant for long time and tend to be filled by biologists with a bit of coding experience, underqualified IT people or, rarely, really competent individuals that want a break from industry.
I know sufficiently many good programmers who would love to do scientific programming (because they love science) and would immediately accept less pay.
The problem, in my opinion, rather lies in the non-monetary work conditions. For example, in Germany it is nearly impossible to get permanent employment contract when working at a scientific institute and doing something remotely related to scientific work. Even worse: you are not even allowed to work more than 6+6 years (before and after doctorate) in a fixed-term employment position at a scientific institute If this time is over, you are not even allowed to take any non-permanent contract at a scientific institute (the infamous/insane Wissenschaftszeitvertragsgesetz (WissZeitVG)).
No programmer (even if (s)he has a great passion for science) will be willing to work under such extraordinarily bad conditions.
Here it is a bit better. Most positions I know of are de facto permanent. On paper they are not, as most labs go through 5-year funding cycles. So there's a tiny chance of loosing funding. It's quite rare for big labs.
Besides, places like MRC have created permanent research assistant positions. Which are actually permanent and put zero pressure on publications.
As you say, connecting with interested and talented programmers is another problem. I feel that Nature Jobs postings, which were already a big leap forward for rusty uni administrators, are not good enough.
I have to ask how you came to this conclusion. Did you have anecdotes from them supporting this or did you just conclude "they are not paying for it therefore software is not worth paying for to them"?
How can you have a ‘non-technical scientist’? All science is inherently technical.
Are you using ‘technical’ as a short-hand for ‘can program’? Stop doing that - programming is not the only technology.
I hate to have to create this conversation fork but I really wish people wouldn't make comments like this. They're so low signal.
The effects are more pronounced depending on what field we're talking about. For example, in physics, I'd imagine most people have at least the fundamentals of programming down, even if software design may be lacking. In the field I work in (Neuroimaging), a lot of PI's are doctors or neuropsychologists, and might barely even know how to use a computer.
I know for a fact it's happened at least once, and led to a scientific controversy that lasted for decades. Unfortunately, I don't remember the specifics; if someone recognizes my vague description please step in with a citation. But one group of scientists published a paper saying a certain dynamic system behaved in a certain way, and a second group published a paper saying it behaved in a different way, and the two results were completely incompatible with each other. Significant public disagreement ensued. One group published their code, and the other group attacked it saying it was poorly written, etc. The second group did not release their code. Decades later, some other scientist at the second institution released the code, and after a code review, it was found that the data was incorrectly initialized. They needed to initialize the particles with "random" initial velocity vectors, but the scientist who wrote the code didn't know how to do it correctly, and wrote an ad-hoc algorithm that gave the initial velocity vectors significant bias along an axis. But the paper was already written, peer reviewed, published, and cited, so even though the paper was wrong, the result was still accepted by (half of) the scientific community. AFAIK the paper was never retracted.
I think most people who have created something will generously bend over backwards to help individuals in the early stages of it's lifecycle. You can see that all the time on github.
The problems come when the project takes off to the point where there isn't enough support for the number of people using it BUT the software isn't mature/popular/fit-enough to be "under the wing" of a larger organization who can afford to pay for it's maintenance and evolution.
Is there a way to bridge the gap between author's-generosity-support and corporate/organizational stewardship? We do have the social networks in place to allow that, they're just focused on different objectives.
Found it: https://arstechnica.com/information-technology/2018/11/hacke...
I don't blame anyone in this scenario because the culture of open source projects and their interplay with enterprise encourages it
It makes me want to give my customer's the source and tell them they can do whatever they want, and then ignore the rest of the community except for high-quality pull requests.
If anyone out there has suggestions of useful resources, I'm all ears!
() Those of us who care about software do try to work on all these issues, but progress is slow.
Question now is: which is better? In one case there is a single "God" in the other it's passed on for generations, while everybody mostly cares about their research and not long term maintainability.
It's in Python. I've avoided any external dependencies, kept inputs to CSV files that can be made from existing Excel sheets, and the code is fairly well commented.
But there used to be two people here who've written at least a line of python in their lives. Now it's just me, and if I leave I have no illusions that it'll be maintained.
Best thing to do is write instructions for whoever will need to run it, and they can hope that they never need it to do anything new.
We develop software in house. I am the single dev on staff, we have a few contractors that we use for legacy ERP system programming/maint./modifications. I work with people who have CS degrees, however, it is still hard for them to understand why I'm spending time on layered architecture instead of using some basic OOP. Thankfully, they trust me and allow me to do what is required.
The BEST solution I see is having more than one full stack dev but then you are paying 2x more. Also, have some kind of standard and review any outsourced work. I picked up a legacy app after an outside contractor and it has been a disaster to work with.
The source code provided was out of date, he could not produce the source code that was running in production. No naming conventions were used, literally everything was just generic names like command1, textbox1, and so on. This could have been easily caught if any competent junior looked at the code. Some methods with tens of if statements, methods that are 1k+ lines long, almost zero OOP.
If a company does have a single dev and cannot afford another they really need to stress maintenance and verify in some way that the dev is capable of producing a maintainable project. Therefore, they maybe should hire someone to help with the hiring process but hiring SEs is difficult even when experienced people are doing the hiring.
[1] https://discourse.julialang.org/t/biostar-handbook-computati...
Even other projects that tried to address this like the Core Infrastructure Initiative seem to have unintended consequences. For example OpenSSL got CII funding, then used some portion of that to relicense as Apache 2 which breaks compat with the more free LibreSSL fork, weakening the overall community.
the problem here is that lots of people want _other_ people to work for free. if you're not being paid, it's a hobby and you don't owe anyone anything. if the science system relies on free labour and refuses to support it, that's a very different conversation and the results are predictable.
This is the most ridiculous statement. I had chance to work in either environment. Having that experience for the quality of the end result I will take scientist (preferably physicist or mathematician) with self taught software development skills over formally trained agile guru any time.
Of course there are exceptions but ...
Indeed, tons of scientists really have no idea about style, testing, etc. They're happy to just write an imperative C/C++/Python program with no docs whatsoever, run it, and be done with it.
As for writing imperative code in C/C++/Whatever other language they seem to choose: nothing is wrong with that.
But that isn't anywhere near the point of the article. It's talking about scientists that did NOT switch to writing software professionally, but continue to write it merely as a tool to accomplish their primary scientific goals.
Sure, for this they would do what is just enough to solve their specific problem. Just as ANY sound business would do.
There's nothing inherently wrong with writing imperative code, and being imperative doesn't imply the code isn't broken up into components or that it's necessarily difficult to maintain.
From a practical standpoint, why invest in future-proofing your code if it's going to be thrown away once the current paper is finished? No need to make the code readable for the next author, no need to document it, no need to make your code non-monolithic because you're never going to build on top of it later.
On top of this, they were never educated and trained as software engineers. If you're lucky, they may have a pure CS background (and some of the worst code I've seen has been written by academic CS people who are explicitly not software engineers), but most likely they come from various academic disciplines that don't teach how to write code.
I used to work at an academia-focused NLP company, and while they did have some well-structured long-lived projects that we used across several projects, there were also large piles of code that you can tell were just intended to be used once and then forgotten about.
This also reminds me of the way Japanese console developers, such as Squaresoft, used to treat their source code in the '90s. Once the game shipped, they would just wipe the source code from their hard drives and their backups to save space. Hey, it's a pre-2006 console game, it's never going to get patched after release, and this was long before the nostalgia bug bit and publishers realized there was money in porting old games to new platforms. Hey, if you believe that code is only ever going to be used once, you are going to treat it as throwaway code, and that includes actually throwing it away at the end. As a result, many later ports of '90s console games were remade from scratch, consist of old ROMs running in an emulator (sometimes romhacked if the game had never been translated before, like Trials of Mana), rebuilt from a third-party port of the game, or if they're really lucky rescued from early beta code stored on an old computer they forgot to wipe (this is how the FF7 PC version got made in fact). There was a Twitter thread recently, and it's absolutely fascinating.
As for what goes for "formal training" in modern colleges. Well I better not go there.
I'm at an institute of computational biology and that is precisely the problem we are tackling right now. We have a lot of clever people doing a lot of clever things, but a large part of effective software development is developing some good habits that those without training have often never heard of (e.g. how to write clean code, defensive programming, etc.)
As one of the more experienced "developers" (a.k.a. self-taught programmers) on our team, I've been writing up what I consider to be the basic principles of good software development for computational scientists. (https://terranostra.one/posts/Principles-of-Software-Develop..., if you're interested.) Later this week, we're going to have a group meeting to discuss what would be a good method of teaching these principles to new members of the institute. Excited to see what will come of that discussion ;-)