Articles like this are a _huge_ deal for that reason. It's an immense delayed recognition for over a decade of work from a lot of folks.
Articles like this are a _huge_ deal for that reason. It's an immense delayed recognition for over a decade of work from a lot of folks.
At a talk of his it lead to a very heated discussion where an older professor accused him of wasting government money on such nonsense.
By the way, those of that opinion are all professors who wanted me on their labs, but I turned them down...
As indeed, I wrote an analysis framework for my data (of a gaseous detector used for axion search) [0] instead of using an existing framework used by my predecessor. However, things are always more complicated than they seem. Many of those not talked about students who rewrite stuff probably have reasons!
In my case the existing framework [1] was a monster that was bent to allow it to work with the kind of data we have in the first place. In my case my detector had several additional features, which fit _even less_ into the existing framework. It would have been a hack and still a significant amount of work to make it work well.
To be fair, when I started this I expected it to be less work than it ended up being. But that's the story of software development.
The advantages now are significant of course. I know the whole codebase. It does exactly what I want. I can extend it easily as I see fit.
That doesn't mean I didn't also partly procrastinate writing software. Far from it. Hell, there was no reason to write a freaking plotting library (a sort of port of ggplot2 for Nim) [3]. But again, this means my thesis will have plots created natively using a TikZ backend while at the same time provide links to Vega-Lite plots for each and every plot in my thesis (which of course will include the data for each plot!).
Finally, the most important point: A university / professor who only pays me for 20h a week does not get to tell me how I do my PhD.
[0]: https://github.com/Vindaar/TimepixAnalysis [1]: https://ilcsoft.desy.de/portal/software_packages/marlintpc/ [2]: https://github.com/Vindaar/ggplotnim
End result so far? I'm quite respected, still one of the leading researchers in my country on my specific topic, but since I don't have a PhD (because of the aforementioned delays, and some grumpy professors actively pushing against me) I'm starting to lose access to grants and programs.
I'd still do it all again, but with a few tweaks here and there, you know hindsight always helping.
In my (albeit limited) experience, software is a pretty common deliverable from a grant, at least in computational biology. This has also been my experience with more alternative funding sources like CZI and DARPA.
Taken more broadly, I think there is a huge disconnect between what academics are paid to do, and what takes most of their time. Review is unpaid. Grants are not dependent on which journal the results go into, but time could be saved by aiming lower. A salary can be payed from a research grant, while the investigator still has to teach.
For a scientist, writing useful software is a good way to get exposure, build a reputation and get citations. It’s an opportunity to do some different kind of problem solving than usual. It’s also a way of understanding how the software really work (which assumptions are built in, which methods are used, and how does it affect the software’s results?). This does help improve the quality of subsequent results.
A grant typically (there are exceptions, of course) lists things that are going to be studied. How the studying is done is typically down to the people doing the work. It certainly isn’t for grumpy old professors who hear a talk at a conference to judge.
I can never understand the arrogance of folks who would say something like this
It’s hard to tell why people behave like that. I presume it’s because many people have a great fear of originality and require social validation.
Ie., we are easily persuaded that something very complex will be very powerful (eg., a smart phone) -- but we intuitively regard something simple (eg., a hammer) as under-powered.
Hard to say how well this actually holds, but I'd guess in both cases we arent really enumerating use-cases in our head, we're just using explanatory complexity as a guide to practical power.
This is probably more extreme in cases where people have a specific notion of complexity in mind, eg., in academic environments where "tool A" is as simple as "tool B" if they use the same theoretical basis.
ie., Tool C is worthwhile if it includes a more complex theory, as therefore it is more powerful.
Python 1.5 was the first non-mentat tier open source interpreter available that didn't get in your way as a scientist, and Numeric/Numpy was the first and still the most elementary piece that made it usable to science and numerics people. Might not have been letters to Nature tier back then, but Nature ain't what it used to be anyhow.
[0] Since a lot of folks have actually forgotten Paul: http://www.pfdubois.com/bio.html
[1] https://en.wikipedia.org/wiki/IDL_(programming_language)
but 'building the underlying infrastructure that tons of people use' is not science. in my department we had to fail a phd student because 90% of his work was just implementing bunch of existing methods as a python library. useful, yes; science, no. wasn't his fault, had a shitty supervisor, but making useful tools is not the same as undertaking scientific research.
What could be useful is openness in used tools and software and a way of getting citation counts for software used. It's nothing more than a table. That way the hotness of publication could start to flow for the underlying tools.
no. scientific research is proposing a useful model of an observable phenomenon. this is what you train for during a phd, at least in natural/life sciences: you learn how to test a hypothesis, not an easy skill.
refactoring code or transforming bunch of C++ into a python library is useful, but it's not science.
> What could be useful is openness in used tools and software and a way of getting citation counts for software used. It's nothing more than a table. That way the hotness of publication could start to flow for the underlying tools.
agreed 100%
For me, our discussion is mainly in where to draw the line around "the process of science". The chair, laptop and coffee machines aren't science. The statistical methods, papers and engineering are. You seem to cut parts of the engineering out, namely the non-novel parts. There's a lot to say for that. But a PhD is proof of apprenticeship as well. I wouldn't grant someone a PhD if all of his work is 'mere retooling'. But in a mainly research papers based PhD-application I wouldn't feel some retooling couldn't be allowed. One could demonstrate scientific craftsmanship in retooling.
if X is e.g. microbiology then it's fair to ask whether (1) some python library proposes something in terms of microbiology, and (2) bunch of biologists should make that decision.
this is why refactoring code is mostly dismissed as 'doing science' by most phd supervisors. sure counts as 'developing skills', which certainly should feature prominently as part of your training, but it cannot be all there is to a project.
If you write code that allows science to be done that couldn't be done otherwise then that is science. As a high profile example, a large amount of specialist software was developed for the LHC to allow it to process all the events coming from the detectors.
It sounds like the refactoring here was not really that useful in the first place.
doing a phd -> training to be a scientist.
I worked on software development tools used directly for LHC as part of an internship.
That experience was of zero use when I tried to apply for a PhD later. It did get me several $BIGN internships though.
Make what you want of this story.
However, there's increasingly a role for folks focused more on the scientific computing and methods side. E.g. "how do we constrain X parameters given Y observations" (yes, I just described inverse theory -- that's deliberate). The science isn't solving the problem, it's figuring out what models to use and what the inverted parameters mean. However, solving the problem correctly requires a lot of rather novel work and is very easy to get wrong.
It's similar to many other research staff positions. It's standard to include the person who operated/designed/etc the instrument you're using as an author on papers. Is it that crazy to include the person who developed the numerical methods and implemented the solution as well? For example, I have quite a few friends that stayed on as staff to run the lab or key pieces of equipment. They have tons of "middle author" publications as a result.
However, numerical methods and computing infrastructure and work is much less frequently recognized. This is a step towards changing that.
no-one is proposing that numpy isn't useful or people developing / maintaining tools aren't doing gods work. they have my endless gratitude and try to donate regularly.
however, phd training in my field -- natural/life sciences -- has a specific remit: you learn how to build and test a hypothesis, from start to end. optimising libraries is emphatically not it. as a scientist you should care whether you have a useful model that explains something about the world. this is orthogonal to how neatly you have implemented your linear algebra in python.
I think journal editors have a responsibility here too in promoting references to software libraries used in the articles they publish. I almost never see these in my field (astrophysics), even though they are readily available and very easy to include.
> for over a decade
Probably two even if NumPy came out in 2005