Array Programming with NumPy
nature.com
nature.com
Articles like this are a _huge_ deal for that reason. It's an immense delayed recognition for over a decade of work from a lot of folks.
I can never understand the arrogance of folks who would say something like this
It’s hard to tell why people behave like that. I presume it’s because many people have a great fear of originality and require social validation.
Ie., we are easily persuaded that something very complex will be very powerful (eg., a smart phone) -- but we intuitively regard something simple (eg., a hammer) as under-powered.
Hard to say how well this actually holds, but I'd guess in both cases we arent really enumerating use-cases in our head, we're just using explanatory complexity as a guide to practical power.
This is probably more extreme in cases where people have a specific notion of complexity in mind, eg., in academic environments where "tool A" is as simple as "tool B" if they use the same theoretical basis.
ie., Tool C is worthwhile if it includes a more complex theory, as therefore it is more powerful.
At a talk of his it lead to a very heated discussion where an older professor accused him of wasting government money on such nonsense.
By the way, those of that opinion are all professors who wanted me on their labs, but I turned them down...
As indeed, I wrote an analysis framework for my data (of a gaseous detector used for axion search) [0] instead of using an existing framework used by my predecessor. However, things are always more complicated than they seem. Many of those not talked about students who rewrite stuff probably have reasons!
In my case the existing framework [1] was a monster that was bent to allow it to work with the kind of data we have in the first place. In my case my detector had several additional features, which fit _even less_ into the existing framework. It would have been a hack and still a significant amount of work to make it work well.
To be fair, when I started this I expected it to be less work than it ended up being. But that's the story of software development.
The advantages now are significant of course. I know the whole codebase. It does exactly what I want. I can extend it easily as I see fit.
That doesn't mean I didn't also partly procrastinate writing software. Far from it. Hell, there was no reason to write a freaking plotting library (a sort of port of ggplot2 for Nim) [3]. But again, this means my thesis will have plots created natively using a TikZ backend while at the same time provide links to Vega-Lite plots for each and every plot in my thesis (which of course will include the data for each plot!).
Finally, the most important point: A university / professor who only pays me for 20h a week does not get to tell me how I do my PhD.
[0]: https://github.com/Vindaar/TimepixAnalysis [1]: https://ilcsoft.desy.de/portal/software_packages/marlintpc/ [2]: https://github.com/Vindaar/ggplotnim
End result so far? I'm quite respected, still one of the leading researchers in my country on my specific topic, but since I don't have a PhD (because of the aforementioned delays, and some grumpy professors actively pushing against me) I'm starting to lose access to grants and programs.
I'd still do it all again, but with a few tweaks here and there, you know hindsight always helping.
In my (albeit limited) experience, software is a pretty common deliverable from a grant, at least in computational biology. This has also been my experience with more alternative funding sources like CZI and DARPA.
Taken more broadly, I think there is a huge disconnect between what academics are paid to do, and what takes most of their time. Review is unpaid. Grants are not dependent on which journal the results go into, but time could be saved by aiming lower. A salary can be payed from a research grant, while the investigator still has to teach.
For a scientist, writing useful software is a good way to get exposure, build a reputation and get citations. It’s an opportunity to do some different kind of problem solving than usual. It’s also a way of understanding how the software really work (which assumptions are built in, which methods are used, and how does it affect the software’s results?). This does help improve the quality of subsequent results.
A grant typically (there are exceptions, of course) lists things that are going to be studied. How the studying is done is typically down to the people doing the work. It certainly isn’t for grumpy old professors who hear a talk at a conference to judge.
> for over a decade
Probably two even if NumPy came out in 2005
but 'building the underlying infrastructure that tons of people use' is not science. in my department we had to fail a phd student because 90% of his work was just implementing bunch of existing methods as a python library. useful, yes; science, no. wasn't his fault, had a shitty supervisor, but making useful tools is not the same as undertaking scientific research.
What could be useful is openness in used tools and software and a way of getting citation counts for software used. It's nothing more than a table. That way the hotness of publication could start to flow for the underlying tools.
no. scientific research is proposing a useful model of an observable phenomenon. this is what you train for during a phd, at least in natural/life sciences: you learn how to test a hypothesis, not an easy skill.
refactoring code or transforming bunch of C++ into a python library is useful, but it's not science.
> What could be useful is openness in used tools and software and a way of getting citation counts for software used. It's nothing more than a table. That way the hotness of publication could start to flow for the underlying tools.
agreed 100%
If you write code that allows science to be done that couldn't be done otherwise then that is science. As a high profile example, a large amount of specialist software was developed for the LHC to allow it to process all the events coming from the detectors.
It sounds like the refactoring here was not really that useful in the first place.
doing a phd -> training to be a scientist.
I worked on software development tools used directly for LHC as part of an internship.
That experience was of zero use when I tried to apply for a PhD later. It did get me several $BIGN internships though.
Make what you want of this story.
For me, our discussion is mainly in where to draw the line around "the process of science". The chair, laptop and coffee machines aren't science. The statistical methods, papers and engineering are. You seem to cut parts of the engineering out, namely the non-novel parts. There's a lot to say for that. But a PhD is proof of apprenticeship as well. I wouldn't grant someone a PhD if all of his work is 'mere retooling'. But in a mainly research papers based PhD-application I wouldn't feel some retooling couldn't be allowed. One could demonstrate scientific craftsmanship in retooling.
if X is e.g. microbiology then it's fair to ask whether (1) some python library proposes something in terms of microbiology, and (2) bunch of biologists should make that decision.
this is why refactoring code is mostly dismissed as 'doing science' by most phd supervisors. sure counts as 'developing skills', which certainly should feature prominently as part of your training, but it cannot be all there is to a project.
no-one is proposing that numpy isn't useful or people developing / maintaining tools aren't doing gods work. they have my endless gratitude and try to donate regularly.
however, phd training in my field -- natural/life sciences -- has a specific remit: you learn how to build and test a hypothesis, from start to end. optimising libraries is emphatically not it. as a scientist you should care whether you have a useful model that explains something about the world. this is orthogonal to how neatly you have implemented your linear algebra in python.
However, there's increasingly a role for folks focused more on the scientific computing and methods side. E.g. "how do we constrain X parameters given Y observations" (yes, I just described inverse theory -- that's deliberate). The science isn't solving the problem, it's figuring out what models to use and what the inverted parameters mean. However, solving the problem correctly requires a lot of rather novel work and is very easy to get wrong.
It's similar to many other research staff positions. It's standard to include the person who operated/designed/etc the instrument you're using as an author on papers. Is it that crazy to include the person who developed the numerical methods and implemented the solution as well? For example, I have quite a few friends that stayed on as staff to run the lab or key pieces of equipment. They have tons of "middle author" publications as a result.
However, numerical methods and computing infrastructure and work is much less frequently recognized. This is a step towards changing that.
Python 1.5 was the first non-mentat tier open source interpreter available that didn't get in your way as a scientist, and Numeric/Numpy was the first and still the most elementary piece that made it usable to science and numerics people. Might not have been letters to Nature tier back then, but Nature ain't what it used to be anyhow.
[0] Since a lot of folks have actually forgotten Paul: http://www.pfdubois.com/bio.html
[1] https://en.wikipedia.org/wiki/IDL_(programming_language)
I think journal editors have a responsibility here too in promoting references to software libraries used in the articles they publish. I almost never see these in my field (astrophysics), even though they are readily available and very easy to include.
"Citing packages in the SciPy ecosystem" lists the existing citations for SciPy, NumPy, scikits, and other -Py things: https://www.scipy.org/citing.html ( source: https://github.com/scipy/scipy.org/blob/master/www/citing.rs... )
A better way to cite requisite software might involve referencing a https://schema.org/SoftwareApplication record in JSON-LD, RDFa, or Microdata; for example: https://news.ycombinator.com/item?id=24489651
But there's as of yet no way to publish JSON-LD, RDFa, or Microdata Linked Data from LaTeX with Computer Modern.
Many labs are gaining access to or creating physical tools that create data analysis over experimental design problems. Biologists transitioning from running gels to detect the existence of a gene to running sequencing or flow cytometry to segment populations, for example. I remember seeing I think an RNA-seq experiment that tracked the full lineage of hematopoietic (blood) stem cells by a professor - my jaw was on the floor at the level of insight.
One remaining step is transitioning many bioinformatics courses from applied tooling to general program design and open-sourcing. I’ve noticed a few labs have done that really well for years, but it is not often found as a field-level skillset.
[1] https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3273988/figure/...
APL from 1966, I believe, is the key lang for array programming.
And statistical languages like S from 1976 come to mind: https://en.wikipedia.org/wiki/S_(programming_language)
At a quick glance, it seems PDL is just a variation on S.
For my understanding, the numpy syntax most closely resembles what would be possible in matlab. And matlab again seems to have roots from Fortran. Thanks to that, young folks nowadays can switch so easily between Fortran and Numpy, the syntax and call structures can easily be made to almost fit to each other.
i think you meant "competition" here :)
(in polish, my native language, it's "konkurencja", but it's a "false friend of the translator"; i'm guessing you're in a similar boat)
Then Travis started Numpy which somehow magically was backward compatible with both numeric and numarray - and managed to get the community united again.
I wonder if github should add a "Review" feature to provide a similar content authoring experience.
There's tons of math and physics blogs that contain useful results that the author wanted to make available but didn't manage to incorporate into a paper. I wonder if there'd be any interest in a sort of GitHub for proofs? It could even use git, since (assuming consistency) isn't math just a DAG anyways (and therefore isomorphic to a neural net, as are all things).
What's missing is the dissemination piece. Somehow people will absolutely refuse to take seriously the job of citing code they use, even when their main result is obtainable by "and then I ran something from scipy/numpy/pytorch/etc."
If you have repo2docker REES dependency scripts (requirements.txt, environment.yml, postInstall,) in your repo, a BinderHub like https://mybinder.org can build and cache a container image and launch a (free) instance in a k8s cloud.
Journals haven't yet integrated with BinderHub.
Putting the suggested citation and DOI URI/URL in your README and cataloging citations in an e.g. wiki page may increase the crucial frequency of citation.
A Linked Data format for presenting well-formed arguments with #StructuredPremises would help to realize the potential of the web as a graph of resources which may satisfy formal inclusion criteria for #LinkedMetaAnalyses.
AFAIU, e.g. Zotero and Mendeley do not crawl and index articles or attempt to parse bibliographic citations from the astounding plethora of citation styles [citationstyles, citationstyles_stylerepo] into a citation graph suitable for representative metrics [zenodo_newmetrics].
bitcoin.org/bitcoin.pdf does not have a DOI, does not have an ORCID [orcid], and is not published in any journal but is indexed by e.g. Google Scholar; though there are apparently multiple records referring to a ScholarlyArticle with the same name and author. Something like "Hell's Angels" (1930)? No DOI, no ORCID, no parseable PDF structure: not indexed.
AFAIU, Google Scholar does not yet index ScholarlyArticle (or SoftwareApplication < CreativeWork) bibliographic metadata. GScholar indexes an older set of bibliographic metadata from HTML <meta> tags and also attempts to parse PDFs. [gscholar_inclusion]
Google Scholar is also not (yet?) integrated with Google Dataset Search (which indexes https://schema.org/Dataset metadata).
FigShare DOIs and Zenodo DOIs are DataCite DOIs [figshare_howtocite, zenodo_principles]; which apparently aren't (yet?) all indexed by Google Scholar [rescience_gscholar].
IIUC, all papers uploaded to https://arxiv.org are indexed by Google Scholar. In order for arxiv-vanity.org [arxiv_vanity] to render a mobile-ready, font-resizeable HTML5 version of a paper uploaded to ArXiV, the PostScript source must be uploaded. Arxiv hosts certain categories of ScholarlyArticles.
JOSS (Journal of Open Source Software) has managed to get articles indexed by Google Scholar [rescience_gscholar]. They publish their costs [joss_costs]: $275 Crossref membership, DOIs: $1/paper:
> Assuming a publication rate of 200 papers per year this works out at ~$4.75 per paper
[citationstyles]: https://citationstyles.org
[citationstyles_stylerepo]: https://github.com/citation-style-language/styles
[gscholar_inclusion]: https://scholar.google.com/intl/en/scholar/inclusion.html#in...
[figshare_howtocite]: https://knowledge.figshare.com/articles/item/how-to-share-ci...
[zenodo_principles]: https://about.zenodo.org/principles/
[zenodo_newmetrics]: https://www.frontiersin.org/articles/10.3389/frma.2017.00013...
[rescience_gscholar]: https://github.com/ReScience/ReScience/issues/38
[arxiv_vanity]: https://www.arxiv-vanity.com/
[joss_costs]: https://joss.theoj.org/about#costs
[orcid]: https://en.wikipedia.org/wiki/ORCID
We'd just need a dedicated search engine, and a way to automatically extract those from papers, to clone and archive repos.
Git uses SHA-1, a hardened version since 2017, and are now doing per-repo upgrades to SHA-256 [0]. Lots of repos are presumably still on SHA-1 (and users on older versions of git).
As of 2020, chosen-prefix attacks against SHA-1 are now practical. [verbatim from 1] But I don't think second preimage attacks are practical yet.
Linus Torvalds argued in 2006 basically that it's irrelevant whether git's hash function is second preimage resistant. Selective quoting:
> remember that the git model is that you should primarily trust only your _own_ repository [2]
> [a malicious] collision is entirely a non-issue: you'll get a "bad" repository that is different from what the attacker intended, but since you'll never actually use his colliding object, it's _literally_ no different from the attacker just not having found a collision at all [2]
All that is just to say: git originally chose its hashes for the above mentioned "git model", thus didn't 100 % care about second preimage resistance. For your suggested search engine, depending on how the database is collected you might not be able to trust "your own repository" (if it's crowdsourced I could register another codebase with the same hash as Linux). A second preimage resistant hash function would be a requirement for the suggested use case.
[0]: https://git-scm.com/docs/hash-function-transition/
Put another way: if I was going to cite numpy, would I cite this? Probably not. Would I cite this paper for any of the more general concepts it covers? Probably not. I'd probably even argue someone shouldn't cite it for that latter reason, as those concepts supercede numpy (and appear in other languages under other names).
This kind of article heralds its adoption to mainstream biology. It's now well known enough to interest biologists in general!
At first I was like, NumPy's fame -- owing to the rise of Python in scientific computing and data science circles -- has gone far beyond that of most things ever published in academic journals, so this hardly seems necessary.
But I see your point: Nature has always been held in high regard among natural scientists, and though most recently minted natural scientists have at least a passing familiarity with Python, the generations of scientists before them probably don't. Nature at least has enough of a cachet to grab their attention.
Plus having a publication in Nature does open doors. Travis Oliphant and some of the already-famous co-authors probably don't need these doors opened, but I'm sure others on that list would benefit.
For most scientists, programming productivity matters the most, and plenty of programs are embarassingly parallel.
For instance it's no trouble at all just launching a single threaded Python program once per sequencing sample, and it plays nicely with the supercomputer queuing system.
That is debatable
For instance, the interface and array/matrix types make vector operations really natural and efficient in python.
MATLAB was created as an interface to LINPACK/EISPACK without having to learn FORTRAN. The importance of this comment is the emphasis on the core fundamentals shared by all the data science platforms rather than the different tradeoffs inherent in each ecosystem.
Numpy approaches the performance you could get with Fortran. Mostly because its core is written in Fortran. What Numpy offers that Fortran never did, though, is leverage. The article mentions, but doesn't really do justice to, the sheer volume of interoperability that Numpy has enabled. It's not just that all these libraries were built on top of Numpy. It's also that their common Numpy substrate makes them all deeply interoperable with each other. And that works both above and below the boundary. You can swap out BLAS and LAPACK for something else - say, CUDA, or a distributed representation - and as long as the replacement also speaks Numpy's language, you can plug it into existing libraries that were originally written against Numpy.
In short: Fortran gets you performance. Numpy gets you that, and also productivity. I would argue that that actually makes Numpy more powerful than what was possible with just Fortran.
In D you get the consistency of a single unified language semantic unlike the impedance mismatched approach that is inherent in Python and Numpy programming combination.
[1]http://blog.mir.dlang.io/glas/benchmark/openblas/2016/09/23/...
Now if you want to use array comprehensions, the first case looks similar: [A[k] for k in range(3,0,-1)] but the second now has to be [A[k] for k in range(2,-1,-1)].
Further, for some reason, array and matrix are different types and one has to convert back and forth between them.
I guess the generic way to write that would be A[bottom:top+1][::-1]. But the blame there goes to Python, not numpy, since the same is true of lists.
https://trends.google.com/trends/explore?date=today%205-y&ge...
That's not a small issue. The ecosystem is probably the reason people choose NumPy over MATLAB, for example. NumPy is not inherently superior to MATLAB, and most academicians that adopted NumPy in the 2000's already had a MATLAB license, so cost was not a concern either.
I do disagree strongly with the opinion that Numpy is no better than MATLAB :). MATLAB has adopted some Numpy features after Numpy came out (broadcasting for example) but Numpy offered some genuine and unique advantages, both technical (broadcasting, no need for a MEX compiler that I have to pay through my nose for, not restricted to weird naming conventions, nature of parameter passing, ...) and legal.
For more pure research and prototyping things both can do, I still think matlab is better though I rarely use it. I just like the idea of being able to easily deploy the code later somehow. Kind of an entrepreneurial feature.
In fact, it's not an issue at all since Julias ecosystem is a superset of that of Python: with PyCall you can use Python libraries and Julia libraries in one program without issues.
using PyCall
np = pyimport("numpy")
res = np.fft.fft(rand(ComplexF64, 10))
You just called numpy fft from Julia.
julia> data = rand(ComplexF64, 1024^2);
# python fft from julia:
julia> res = @btime np.fft.fft(data);
78.613 ms (39 allocations: 16.00 MiB)
# python fft in ipython:In [11]: %timeit res = np.fft.fft(data)
89.3 ms ± 1.65 ms per loop (mean ± std. dev. of 7 runs, 10 loops each)
As expected, julia has its own fft package (based on FFTW):
julia> res = @btime fft(data);
61.540 ms (33 allocations: 16.00 MiB)Python packages are either interfacing external libraries (something that is much easier to do in Julia) or if they are pure python, badly designed and buggy.
(and package management in python is broken beyond repair)