Fortran vs Python: The counter-intuitive rise of Python in scientific computing (2020)
fortran-lang.discourse.group
fortran-lang.discourse.group
In that regard, I'm surprised Nim hasn't taken off for scientific computing. It has a similar syntax to Python with good Python iterop (eg Nimpy), but is competitive with FORTRAN in both performance and bit twiddling. I would have thought it'd be an easier move to Nim than to FORTRAN (or Rust/C/C++). Does anyone working in SciComp have any input on this - is it just a lack of exposure/PR, or something else?
Combined with the ease of displaying results ala Matlab but much less of the jank, and you have an ideal general purpose sci comp environment
It's tempting to get lazy and just use a for-loop to iterate over an array sometimes and that will absolutely kill your performance.
If I were to start again today, I think I'd give Julia a look, though.
End users are using python - the advantage of modern computing is that whatever happens afterwards is irrelevant.
Python's success starts with academia's movement of replacing Matlab with free software, namely numpy/scipy.
That makes any sort of experimentation a really tough sell.
As a rule, I have found scientific computing (at least in astronomy, where I work) to be very socially pressured. Technical advantages are not nearly as important as social ones for language or library choice.
Change does happen, but extremely slowly. I am not exaggerating when I say that even in grant applications to the NSF as recently as 2020, using Python was considered a risky use of unproven technology that needed justification.
So, yeah, Nim is going to need a good 30 years before it could plausibly get much use.
My guess is that grad students are swamped and are looking for the shortest path to getting an interesting result, and that is most likely done with a tool they already somewhat know.
The question for Nim, like many other new products, is: why is it worth the onboarding cost?
It’s good engineering and good management but research shouldn’t really care.
Many grad students forget that their main purpose is to generate research results and to publish papers that advance the field, not to play around with cool programming languages (unless their research is about coding).
Here's a bunch of mistakes I made in grad school which unnecessarily lengthened my time in the program (and nearly made me run out of stipend money):
* Started out in Ruby because I liked the language, but my research involved writing numerical codes, and at the time there just wasn't much support for it so I ended up wasting a lot of time writing wrappers etc. There was already an ecosystem of tools I could use in MATLAB and Python but nooo, I wanted to use Ruby. This ended up slowing me down. I eventually gave in to MATLAB and Python and boy everything just became a lot easier.
* Using an PowerPC-based iBook instead of an Intel Linux machine. Mac OS X is a BSD (plus I was using a PPCarch) and Brew didn't exist back then, so I ended up troubleshooting a lot of compile errors and tiny incompatibilities because I liked being seen to be using a Mac. When I eventually moved to Linux on Intel, things became so much easier. I could compile stuff without any breakages in the one pass.
I also knew a guy who used Julia in grad school because it was the hot new performant thing when all the tooling was in Python. I think he spent a lot of time rejigging his tooling and working around stuff.
Ah the follies of youth. If only someone had pulled me aside to tell me to work backwards from what I really needed to achieve (3 papers for a Ph.D.) and to play around with cool tech in my spare time.
I guess the equivalent of this today is a grad student in deep learning wanting to use Rust (fast! memory-safe! cool!) even though all the tooling is in Python.
I like Nim a lot. And I know that it'll scratch a necessary itch if I'm working with scientists. I also know that it's too much to ask that the scientists just buckle down and learn Rust or something like that.
But as someone who is not afraid of Rust but is learning Nim because of its applicability to the crowd that I want to help... The vibrancy of the Rust community is really tempting me away from this plan.
I've really enjoyed the Nim community also. I even contributed some code into the standard library (a first) and was surprised at how easy they made it.
But I have also written issues against Nim libraries which have gone unanswered for months. Meanwhile, certain rust projects (helix, wezterm, nushell) just have a momentum that only Nim itself can match.
Python benefitted from there being no nearby neighbors which resembled it (so far as I'm aware). If you needed something like python, you needed python.
Rust and Go and Zig are not for scientists, but they're getting developer attention that Nim would get if they didn't exist. Also, Julia is there to absorb some of the scientist attention. It's a Tower of Babel problem.
I can't say why the scientists aren't flocking to Nim, but as someone who wants to support them wherever they go, this is why I'm uncertain if Nim is the right call. But when I stop and think about it, I can't see a better call either.
Because most scientists are only using programming as a tool and don't care one bit about it beyond what they need it to do. They don't go looking for new tools all the time, they just ask their supervisor or colleague and then by default/network effects you get Python, Fortran, or C"++". You need a killer argument to convince them to do anything new. To most of them suggesting a new language is like suggesting to use a hammer of a different color to a smith - pointless. With enough time and effort you can certainly convince people, but even then it's hard. It took me years to convince even just one person to use matplotlib instead of gnuplot when I was working in academia. You can obviously put that on my lack of social skills, but still.
It's sad, as I feel nim would be easier to maintain compared to a typical c or R codebase written by a biologist, but that's what's expected.
I for instance just moved to a company where the data stack is basically OracleSQL and R. And I dislike both. But as _Wintermute pointed out, a whole company / department won't change their entire tech stack just to please one person.
In addition to Nim, D programming is also Phytonic due to its GC by default approach and it is a very attractive Fortran alternative for HPC, numerical computation, bit twiddling, etc. D support for C is excellent and the latest D compiler can compile C codes natively, and it is in GCC eco-system similar to Fortran. Heck, D native numerical library GLAS is already faster than OpenBLAS and Eigen seven years ago [1]. In term of compilation speed D is second to none [2].
[1] Numeric age for D: Mir GLAS is faster than OpenBLAS and Eigen:
http://blog.mir.dlang.io/glas/benchmark/openblas/2016/09/23/...
[2]C++ Compilation Speed:
What seems to have bootstrapped the success of Python for ML and scientific use was early adoption by people in these communities who were not hard core programmers, and found it easy to get started with. Once SciPy and NumPy were available, and NumPy became used for ML, then the momentum of ML helped further accelerate the adoption of Python.
Why it's deemed unsuitable for large, complex, multi-person projects is that enterprise types know only the byzantine OOP mess. And when all you have is OOP, everything is a FactoryFactoryFactory and everything else is "unmanageable".
I disagree that the availability of good libraries for Python is because of the language - especially if we're talking about scientific and ML libraries. In many of these cases the Python libraries are just pass-thrus to the underlying libraries written in C, which was chosen because of it's performance and suitability to the task.
How so?
I guess if you're doing team-based Python development then everyone is going to be forced to be sensitive to indent, and maybe use a Python-aware editor, but often in a large non-Python project there are many different editors and indent configurations being used and the code structure is unreadable - but at least you can use a language aware formatter or indent tool to recover it. In Python any sort of loss of proper indenting would obviously change the functionality of the code. Maybe this never happens ?
Have you ever worked in a large, dynamically typed codebase written by other people?
Similarly, I've worked on very large, multi-team projects where the language was statically typed, compiled, etc. and they've been total disasters, and others that have been a huge success.
I've yet to see any language that can fully negate the power of sloppy, undisciplined programmers. Those programmers are like water and always find a way.
Sorry, but that makes zero sense.
Dynamic typing is a language design choice where one is trading off automated error detection for faster development. The larger and more complex a project becomes, the more moving parts and interfaces (APIs) it has, and the more potential there is for API errors. Choosing NOT to prioritize automated (compile time) error detection is NOT something that scales up "very, very well".
It's not just about size of project, but also about expected project lifetime. For a one-time use script, or experiment, or thesis project, then maybe development time is the prime concern. You're going to hack it to get it working ASAP, and don't care what happens later (there is no later).
However, for a corporate project that will be in production use for years, then initial coding effort is only a small part of the lifetime project costs, and optimizing this at the expense of ease of maintenance is a poor decision. Years down the road the project will have had staff turnover, developers will have forgotten the details of all the code, the original design has probably been compromised due to feature creep, too many hands, etc, etc. At this stage you want all the help you can get so that people can still maintain the code, and the larger and more complex the project, the more so this will be. Again, the small/throwaway project language choice trade off of optimize development time at the cost of ease of bug detection will NOT scale up "very, very well".
If you anticipate sloppy undisciplined programmers working on a project, then all the more reason to give them less weapons with which to shoot themselves in the foot.
Well, I hate to break it to you, but I've been using Python in large production environments for nearly a quarter of a century now and either I'm just extraordinarily lucky, or it does in fact make quite a bit of sense and really does work quite well. :)
Here's one example: we once had an enterprise Java system of roughly 1 million lines of code. For a number of reasons the whole thing got replaced by a port of it to Python, and the Python version weighed in at just over 100kLOC. Disregarding all of the other benefits (which were many), there were meaningful advantages to maintaining a codebase 1/10th the size.
> Dynamic typing is a language design choice where one is trading off automated error detection for faster development.
No, not at all. First of all, if you are dealing with any sort of production code, you have to be investing in good testing. The idea of dynamic languages hiding bugs that don't show up into production is mostly a bogeyman to cover inadequate testing. We've also found that languages like Python tend to encourage a style of iterative development that lends itself to each section of code being very well tested as it gets written, so in the end it's easy to wind up with code that is both tested more during the dev process but then also well-tested due to the automated tests that all software should end up having anyway. I mean, all software gets tested, it's just a question of whether you do it or your customers do it.
(anecdotally we've seen evidence that people come to over-rely on compile time error checking in lieu of good iterative testing, and that's an interesting topic in and of itself, but beyond the scope of this discussion and IMO falls under the 'sloppy developer' umbrella anyway)
> The larger and more complex a project becomes, the more moving parts and interfaces (APIs) it has
The implicit argument here is based mainly on the assumption of lots of (even exponential) growth in coupling as a project grows, and the reality is that, the level of coupling tends to go in the opposite direction as projects grow large, unless it's simply poorly architected. For example, instead of intercommunication between modules, you're dealing with intercommunication between entire systems - the points of contact between systems tends to be over very well-defined interfaces and are small in number relative to the amount of communication internally between modules.
> initial coding effort is only a small part of the lifetime project costs, and optimizing this at the expense of ease of maintenance is a poor decision
Speed of initial implementation is certainly one benefit, but it actually pays dividends over the lifetime of the project. Code is read more than it is written, so having fewer lines of code, in a language that is easy to read, in a language that removes so much of the noise associated with more verbose languages, is always helpful. It helps you implement stuff to begin with, it helps you add new features later, and it helps you fix bugs.
There are also lots of second-order advantages too, such as making it easier to bring new hires up to speed and needing fewer people both initially and long term. It's about a lot more than just dynamic typing though; a higher level language just provides a ton of benefits that are often worth the tradeoffs. There are so many parallels to e.g. when we moved from assembly to C.
> If you anticipate sloppy undisciplined programmers working on a project
That point was just that so many knocks against certain things are actually covering the problem of sloppy developers. Rather than just take them for a given, a better approach is to set them on a path to improvement if they are willing, or let them go if they are not.
> What seems to have bootstrapped the success of Python for ML and scientific use was early adoption by people in these communities who were not hard core programmers
What if these people (non-hard-core programmers) were attracted to the language itself because it is almost pseudo-like? So it becomes a gift that keeps on giving. Attract domain experts and you get more batteries attached for your project.
> hard core programmers
What if these people are 'hard-core' in their specific domain, but not 'hard-core' in whatever hardware architecture carries the day due to historical mishaps and marketing trends of the day?
Python is a popular user interface for scientific computing, data science or high-performance computing.
Python is the default language in which people express their scientific computations. It may execute C code in the end, but so does any language that ever executes a system call.
Little fun fact: Numpy doesn't even come with an efficient, blocked matmul procedure. It has to be linked against a BLAS implement to really provide any decent performance. This also explains why Numpy performance can vary from distribution to distribution. Anaconda ships it with a different BLAS than Pypi.
That was all, it could have been Tcl instead.
I'd bet that their python usage is still mostly as a REPL to ROOT (which by the way has its own REPL), so no numpy, maybe little pandas, no matplotlib.
However, if I didn't know how things work underneath I'd be a little uneasy. You can always profile after the fact but it helps knowing how to avoid inefficient approaches.
And of course it saves on the insane licensing costs since Mathematica is no longer required in all student software packages (MATLAB still is afaik).
I still maintain that Python, when evaluated strictly on its merits as a programming language, is the most ass of the "scripting" bunch, but its ecosystem is such that it more than makes up for the difference and I always end up using it for side projects or personal stuff.
The "free" thing is important, because programming tools have all migrated to the open source model. No professional coder is willing to work on paid tools today. If scientists have a basic awareness of their career options, they will borrow tools from the software development world, and not from the engineering world. Those tools are all free.
Also, free software changes how you use it. I install a complete Python toolchain (up to and including Jupyter Lab) on every computer that I touch: In the office, the labs, and at home. This allows me to truly use Python as my brain. No software budget is lavish enough to pay for as many "licenses" as I use in a day.
Python, first appearance: 20 February 1991; 32 years ago [1]
Rust, first appearance: May 15, 2015; 8 years ago [2]
Julia, first appearance: 2012; 12 years ago [3]
Go, first appearance: November 10, 2009; 14 years ago [4]
Java, first appearance: May 23, 1995; 28 years ago [5]
Javascript, first appearance: December 4, 1995; 28 years ago[6]
[1] https://en.wikipedia.org/wiki/Python_(programming_language)
[2] https://en.wikipedia.org/wiki/Rust_(programming_language)
[3] https://en.wikipedia.org/wiki/Julia_(programming_language)
[4] https://en.wikipedia.org/wiki/Go_(programming_language)
[5] https://en.wikipedia.org/wiki/Java_(programming_language)
[6] https://en.wikipedia.org/wiki/Javascript_(programming_langua...
Please read what I said again, more carefully, starting from the first word, and try not to be so contrary. I was referring to the fact that the shiny and new tooling will be what attracts users to a given platform / language / ecosystem, regardless of the latter's age. Other participants on this thread apparently seem to have understood this.
Happy new year.
Jetbrains has support because python is popular, and not the other way around.
I was also expressing that Python is an old, almost ancient language, and yet it more used than languages that, by your reasoning, had better "tooling" in the sense that you are using it. Making the point that it has nothing to do with the "age" of the language, the fact that it is perceived as new, or having better tools.
Python isn't popular because of its tooling, but despite it.
It is popular because it makes it easy to leverage the ecosystem and get things done.
It seems that other people understood as much.
I'm a big Fortran now for anything fast. If I was doing EDA or any one-of data science stuff I'd be more likely to use python though (or coreutils and gnuplot depending on the circumstances).
That, and how easy it is to install and import things.
Newer ML courses use Python because of it's subsequent adoption by PyTorch.
We have a bunch of people programming, most of them scientists. Even if Python is poorly suited for us, it’s pretty much the only thing everyone can work with.
Nowaday, it seems that at least physics particle community looks enthusiastic regarding Julia development.
*: it mimics most of core python features such as no strict typing, data structures that can store different types, on-the-fly coding, graphic interface for representation, no compilation needed.
I bet that plays some role too in its popularity in the scientific community, which has many young anxious grad students/postdocs looking to ensure they are employable.
Languages are mostly funglible, coding culutre is not.
Contents:
- “Python, a slower language”
- “more and more”
- “time-critical scientific computations”
- Is Python ever better suited?
- User story : same author, two languages
- Speed vs agility
- Takeaway
For i in range(100):
mask = lib.abs2(z) < 4
subset = z[mask]
subc = c[mask]
subcounts = 0
for j in range(100):
subcounts = subcounts + lib.abs2 (subset )< 4
subz = subz**2 + subc
Z[mask] = subset
Counts[mask] += subcounts
Was that code using numpy, tensorflow, torch, arrayfire, or some proprietary amd gpu lib? It’s hard to say!Try the same with eigen vs arrayfire in C++ or math.js vs tensorflow.js and you will have to do a hell of a lot more than change the value of lib
c = lib.arange(-2, 2, .01)
c += 1j * c.transpose()
at the beginning, and appear to have eaten some downvotes as a result lolIf you care about performance you should care about how vectorization is implemented. Python makes vectorization look like magic, but scientists shouldnt do magic. In Steel Bank Common Lisp I can implement SIMD procedures in a straightforward manner. The language is more high level and more low level (yes) than python, and much much faster.