The skills gap for Fortran looms large in HPC
nextplatform.com
nextplatform.com
In my effort to try to use Python instead in my field, I fell down countless rabbit holes along the way (like being for a year or so one of the core contributors to Cython, in an effort to make "a better Fortran").
In the end though -- I decided it was time to just Get It Done, and just did what my supervisor had patiently and subtly hinted for years. Enough rabbit holes, just get it done in Fortran, and hand in the thesis...
I think what I learned most about (having an interest in compilers and programming languages) is the ignorance the rest of the programming language community, and computer science, has about the kind of things you are actually interested in when doing HPC. What kind of language features you care about and what idioms make sense.
Learning Fortran is super simple.. that is perhaps part of the problem that also people who are barely programmers can write (bad) programs in it. So what this must mean is simply a lack of programmers, full stop, being attracted to the field.
And...C++ is truly the most wrong language for this space there could be.
So if you envision doing programming like that outside HPC after your PhD, C++ might be worth the investment.
But of course a lot comes down to what language other people in your group and your particular subfield, and what your supervisors are using. Doing a PhD is hard enough without being the odd guy out who's using a language nobody else in your community is using.
Some older folks used Perl in places where others would've just written bash; mostly scripting invocations of actual number-crunching programs. Ruby would've raised some eyebrows for sure!
I chose Python over perl about 25 years ago and never would have considered ruby or php (both seemed squarely aimed at web developers not scientists). I believe Numeric Python was released around that time, and it was a revelation- especially to matlab users, who recognized the syntax and behavior.
Python has a low barrier for entry and can be useful immediately to newer programmers. It is also fairly well rounded.
That being said, I am not partial to python as OP was with fortran, and steer around it regularly.
What? Why???
I switched from Python to C++ because Cython, Numba, etc. just weren't cutting it for my CPU-intensive research needs (program synthesis), and I've never looked back.
The kind of C++ written for HPC tends to be much simpler than the kind of C++ used to write e.g. database engines. The complexity of C++ is not that onerous in context and Fortran isn't exactly an exemplar of ease of use.
When C++ one day goes out of fashion, like Fortran already has, this will turn out to be bad news. Fortran is not a big language, and any programmer can learn Fortran in a matter of weeks. But if you'd need to teach people C++ just so that they'd be able to maintain a legacy code base, that will be considerably more difficult than with Fortran.
Point being, the premise of your argument isn't granted. Is C++ going away soon? Or is it holding steady and even growing? How can someone objectively know either way?
But we keep expecting languages like COBOL and MUMPS to die, and they keep living on.
I think there’s a almost 100% chance that 1000 years from now, when all that’s left of human civilization is a bunch of bots trading hustle culture spam and culture war arguments on a zombie internet, that the whole thing will still run on some mission critical code written in C. Buggy, unsafe, but mission critical and hard to replace.
And of course there will be heated bot arguments about how it should all be scrapped and rewritten in Rust++.
You can already give GPT a decent-sized chunk of C code and ask it to rewrite it in, say, Rust, and it works with some rate of success.
For the last 50 years, this was impossible to do in a generic way. In the last 6 months, we managed to move the needle from impossible to possible. Give it another 12000 months, and what are the chances this would not move towards "trivial"?
The problem with changing a mission-critical system is risk, and cost. Maybe omniscient AIs in 3023 won't have this issue, but my ghost will be entirely unsurprised if there are still mission critical systems written in C laying around because whatever benefits are rewrite would give are dwarfed by the risk, perceived or real, of what happens if the rewrite isn't 100% perfect or doesn't 100% match the current system.
No. It's not going away anytime soon, but you can look at high-performance languages that are gaining in popularity, and that may point to change in the future. TIOBE isn't perfect, but it's a good indicator of interest, and Rust and Go seem to continue to gain popularity.
Interestingly, and related to this article, Fortran has moved up from #31 to #20 on TIOBE's index, which really does speak to the importance of a math-optimized, high performance, compiled language. Another interesting change is MATLAB moving up from #20 to #14.
> Most popular technologies This year, we're comparing the popular technologies across three different groups: All respondents, Professional Developers, and those that are learning to code.
Most popular technologies > Programming, scripting, and markup languages
The Top 500 Green500 (and the TechEmpower Web Framework benchmarks) are also great resources for estimating what people did this past year; what are the "BigE" of our models in terms of water, kwh of [directly or PPA-offset] sourced clean energy.
Also, neural ODEs are already taking traction. It won't be hard to replace lots of obscure spaghetti code that deal with edge cases in models with maybe inscrutable neural networks that deliver the same results. Better tools are already developing to deal with interpretability of NN, as higher-level blocks become standard abstractions.
Which the article does admit:
> The skills issue with Fortran is apparently not just about learning Fortran, but more about being associated with Fortran and all of the legacy baggage that has given its vintage and its low marketability going forward.
This quote from the article rings true to me:
>First, the lack of a standard library, a common resource in modern programming languages, makes mundane general-purpose programming tasks difficult. Second, building and distributing Fortran software has been relatively difficult, especially for newcomers to the language. Third, Fortran does not have a community maintained compiler like Python, Rust or Julia has, that can be used for prototyping new features and is used by the community as a basis for writing tools related to the language. Finally, Fortran has not had a prominent dedicated website – an essential element for new users to discover Fortran, learn about it, and get help from other Fortran programmers.
FORTRAN IV/77 has some statements that IIRC correctly have or will be removed from the newer fortran and that I think no other language has. Things like computational gotos, and a few others I have forgotten. Those could cause a bit of confusion with people looking at this old code on HPC.
Unlike COBOL, at least moving from FORTRAN you will not need to deal with rounding issues on floats.
Right, importantly GOTO. The question I've puzzled over for some while is how much legacy (unmodified) FORTRAN IV/77 applications/libraries still exist and are in regular use—and whether porting this old code to later Fortran or other languages remains a significant problem.
My interest stems from the fact my first language was FORTRAN IV.
I'm not so sure that will continue to be the case. Fortran has moved up substantially in TIOBE (to #20), so there's clearly growing interest.
My research project required use of the university mainframe for the big calculations, the 1000x1000 Markov matrices I was working on. I taught myself how to work on out-of-core code then as well, as well as exploiting symmetries inherent in my models.
My Ph.D. code was all written in Fortran, with a little Perl thrown in for good luck. 30+ years later, all of it works without a problem on my laptop/desktop machines.
I learned C in 1996 and C++ in 1997. The C code I wrote still works, though the C++ code needs special compilation options to be able to work.
I've not used Fortran in maybe 15 years now. I'm not up on the modern bits, I stopped using it professionally as F90 was becoming a thing.
I taught graduate HPC programming courses at my alma mater in the CS department, and gave the choice of Fortran or C to the kids. They blanched at the thought of using a nearly (at the time) 60 year old language. They correctly reasoned that learning the language would not help their employability.
Today, I use Julia for heavy computation. Python as a glue language, and C++ when required. I read the report, and I disagreed with the conclusions for a number of reasons, but what it comes down to is that many research groups in Physics/Chem still use Fortran, and aren't about to do the port to C/C++ due to funding. You can't get funding for this conversion. So codes will go on, students will learn what they need, and profs will keep publishing papers.
Moreover, this isn't the first, second, etc. time that people have predicted the death of Fortran. I heard this in 1990s as I was finishing up my Ph.D in theoretical/computational physics.
In 20 years or so, after I have (hopefully) retired, I'd bet that Fortran is still in widespread use, with Python following Java down the path to legacy. And the newer HPC language(s) will be quite happy to talk to Fortran.
I could be wrong, and I'm ok with it. But I doubt it.
If funding's a problem across many fields then wouldn't it perhaps make better sense to just accept Fortran as the most appropriate language for certain well defined application types/fields and concentrate on improving the language, that is fixing its actual or perceived limitations in current environments?
It would seem more efficient to centralize effort to improving just the language and its tools than to convert or rewrite a whole world of disparate applications that have been developed over many decades. After all, take English for instance, at any point in history, say 1600, its grammar and vocabulary were appropriate for the time. Moving on 400+ years till now we didn't chuck English away and replace it with a seemingly better language such Esperanto but progressively updated it to modern requirements.
I have to admit to some bias here in that Fortran was my first language so it acted as the template for others. That said, I am not convinced that the enormous plethora of different languages that have flooded programming in recent decades has benefited computing and CS to the extent that perhaps it ought to have as there has been a great deal of unnecessary duplication and overlap that's led to wasted human effort—the need to learn many different languages, lack of uniformity etc.
The large number of languages, lack of agreed consensus/standards—the most appropriate language for a given class of applications, etc.—has also led to language ghettoization where programmers swear by one language they've become familiar with and continue to use for jobs where another would be more appropriate. And one can't blame them for not wanting to learn a new language seemingly every other year.
It seems to me that rationalizing and simplifying the language problem ought to be a high priority for CS. Given its long and mature history, its entrenched position in certain fields and its proven suitability for math-intensive work, that process could begin with Fortran as it would likely be the least disruptive.
Yes it would. It would make far more sense than writing whitepapers extolling the risks of staying with the language, for example.
> It would seem more efficient to centralize effort to improving just the language and its tools than to convert or rewrite a whole world of disparate applications that have been developed over many decades.
There are ISO committees dedicated to improving Fortran[1].
...
> It seems to me that rationalizing and simplifying the language problem ought to be a high priority for CS.
It is not. This would relegate CS to a different role at a university, more of a tool building and improvement (e.g. engineering) than a "science". Moreover, there is no real money to be made, or reputation to be created by improving a tool. Especially one that has been in use so long.
CS loves to follow/lead with the new shiny thing. This is how the profs get grants. Show their value to the community. Get their students hired and starting companies. Or taking a leave from university, and going to work as chief scientist of AI at large global companies. (cough cough)
Most of these researchers would prefer to show their value and the value of their thoughts/work, by creating new and shiny things in languages or new languages. Yes, this is cynical. I've seen it first hand. I've watched fads/trends wax and wane in CS for a while now. Often times being unaware that much work was being repeated.
The real problem is that no company wants to train and they all expect experts on day one. Even if one company was willing to train, the employee would be SoL when searching for a new job since no other company wants to train.
I'm starting to think that replacing a bachelor requirement for most dev jobs with an associate degree followed by an apprenticeship would be better and cheaper. Perhaps an industry-wide union setting career development paths would be (marginally) better than what we have today, at least for people just starting out or switching jobs.
I was looking into applying to a job at a national lab. Not anything crazy, no HPC or scientific programming, just back office enterprise crap, but they seemed so intent on wanting a degree. I’m not sure places like that in particular that are primarily seen a “scientific” organizations will be willing to ignore credentialism.
Coding is the easy part of software engineering.
With that said, there's obviously a gap between folks that are mainly just programming vs. folks that are engineering systems. And both personas have been lumped in together. The fundamentals that a good CS, math, and stats education teach you are invaluable when you actually have to build things that scale or work on difficult problems.
I agree. What I'm saying is a 2 year program would satisfy that. Basically 2 years of college is for learning non-domain related junk (health class from HS, music history, English proficiency from HS, etc).
I mean the engineer and programmers get lumped together because there is no real separation today. Put "programmers" on a project designed by "engineers" is a disaster. Documentation and communication always leaves something to be desired. Not to mention everyone calls themselves an engineer. Also I work in finance IT.
Some languages on a CV imply extensive knowledge of complex libraries that aren't technically part of the language but considered necessary. Think JavaScript.
* Universities are teaching Python more than ever. When I started my undergrad in Physics everyone took C++, by the time I graduated it was all Python.
* Academia as a career track is not stable so people want to write code in something that they might have a hope of getting a job in afterwards. C/C++ are much more applicable in the wider world so anyone with one eye on the exit is likely to prefer one of those…
* Fortran itself is fine and even improving with fpm and the ecosystem changes happening, but the big Fortran scientific code bases are not much fun to work on.
* There’s no performance benefit to sticking with Fortran. You can make code fast in any language. The main argument that Fortran is faster than C/C++ is to do with aliasing which anyone writing performance critical code will know how to deal with anyway. Performance portability codes like Kokkos that let you run on OpenMP/MPI/CUDA but write code once are pretty much all C++ too.
* A few years ago, CUDA could only be used in Fortran with the PGI compiler which was commercial. This made it less attractive vs C/C++. This has improved now but it was a major factor for a few years. Even now I think they’re not at feature parity although it’s been a long time since I looked.
Nonsense, any physics PhD student could write meaningful Fortran code over a weekend that's significantly faster then all but the most optimized C code. And even then there might still be a small gap.
It's a huge difference in skill-required effort and hours-spent effort, maybe as high as 100 to 1.
Ignoring that Fortran should be easier for a compiler to optimize than C, this is more of a measure of how much performance "fast" code leaves on the table. You can get very good speedups by just paying attention to the problem, but short of reinventing the universe there's only so much you can do.
Ab initio just being a physics PhD doesn't magically give you skills of optimization and (say) knowing how caches work. I know these skills do often come together but in my experience for example every single academic "programmer" I have ever met has been basically lay - fumbling around in the dark, no taste, no idea how the machine works, no idea how to write code that isn't wrong, etc.
This is something most programmers have not even heard about...
This means simple Fortran programs have more room for parallelizing optimizing compilers than C programs (without restrict annotations).
There is a point about making it as easy and idiomatic to write performant code. If a naïve implementation performs as well (or almost as well) as a more difficult to understand explicitly optimized one, the former one tends to win. Easily usable libraries turn that on its head (as we see with Python and its numerics tools on desktops - not sure how adoption in HPC is).
There is certainly truth in that and certainly Fortran has an advantage here vs. C++ due to the aliasing rules and generally lack of subtle performance footguns in the usage of the language. But TBH, I think a lot of the reasons why Fortran is perceived to be fast is that out of the box it doesn't come with anything like a data structures and algorithms library, and domain experts who don't have a strong CS background are either not aware of other data structures or can't be bothered to implement them, so they code like the array is the only data structure known to man. And CPU's love it.
What sort of architecture and libraries are you making use of where this is the case?
In terms of column vs row major order - the underlying implementation will have its own default way of doing it (depends on the implementation) but you can if necessary recast matrix operations in such a way as to avoid actually doing transposes in the interface layer with the other language. so there’s not really a performance penalty as no additional work is done. CuBLAS and MKL’s BLAS implementation even let you pass parameters in about your input data’s memory layout.
Literally incorrect, to this day.
Interestingly, Lisp and Fortran were both initially implemented on the same vacuum-tube based computer, the IBM 704: https://en.wikipedia.org/wiki/IBM_704#Landmarks
A Python/Fortran combination would have made the development of high-performing HPC type applications more accessible to scientists, more fun, and would have given Fortran a more secure future.
The decisions or non-decisions, the culture and approach to community etc of key stakeholders of language ecosystems do have implications.
SciPy makes heavy use of Fortran. Numpy uses Fortran (but it's optional). f2py compiles fortran into python modules with little fuss. Python's ctypes module can directly call fortran libraries.
But to say Python should have been implemented in Fortran rather than C to enable Fortran interop by default -- that doesn't make sense. Python was designed as a scripting language for C-based operating systems, so it needed to be able to use pre-build unix libraries written in C, and it also needed to be embeddable in C programs. At the time neither of these things was easy via Fortran (cpython predates the Fortran 90 standard).
Edit: In case you're confused, cpython is the de-facto standard implementation of Python. Cython is something entirely different (it compiles Python to C). Here's a good summary of Python's history: https://www.geeksforgeeks.org/history-of-python/.
From a Python implementor standpoint, you have to implement Python's interpreter in the language you want to do the tight loops in. So for example, Jython was a reimplementation of Python on the JVM to make it possible to easily use Java for the tight loops. IronPython was the same, but for .NET.
You say it's cultural, but it's all about ergonomics and backwards compatibility. Making it easy to write tight loops in Fortran means changing the FFI API to be fortran friendly, which it is not.
Just try writing your MPI code with a combination of Fortran and Python. I did it circa 2005 and it was a technical nightmare.
Python’s focus on C-extensibility predates its explosion in science and is a “particular reason” for this.
Its not an abstract, perfect-world, ivory tower reason, but most real reasons aren’t.
Code has an almost surreal longevity.
But, then, imagine an HPC parallel sysplex of a thousand five-drawer z16's filled to the gills with Nvidia GPUs for the number crunching.
As a Clojure lover, this hurts.
How can we not feed Fortran source in and receive Language XYZ as output? Perhaps the new wave of ChatGTP will write this and save society.
I guess there's not enough commercial demand, and any hobbyist who's good enough at Fortran probably isn't interested in doing it.
edit: apparently f2py (Fortran to Python) is a thing.
It’s like in professional software, ignoring that we’re generally too willing to decide on a rewrite, one of the main benefits of a rewrite is your current team builds expertise on all the logic or behavior that predates them (and usually uncovers hidden knowledge like implicit behavior, subtle bugs, the code being ugly because it needed to be to solve a bug).
A typo ? I thought FORTRAN was a bit older than COBOL :)
Anyway, i very nice article.
What I've started doing recently is just pasting into ChatGPT and having it explain the code to me. It actually works great.
How does an AI know to surface those concerns?
For the moment nothing, I am not really surprised. It is quite a lot of you know someone who... kind of niche. To port or upgrade some scientific code from Fortran to a new language, you need to have some minimal, some times extensive, domain knowledge, Fortran knowledge and target language knowledge.
It is hard to find the right persons, with not that many persons, you do not want to start a big migration and lose the head developer on the way, etc. This is why if you read in the report and the forums, what people want is to "reboot" the community/ecosystem and put new life in it. But this is also hard if the ecosystem starts to fossilize.
I love and really enjoy working in Fortran, but as a paradox, because of the mentioned situation, it is not that easy to find Fortran work.
An AI that uses LLMs but is more than just an LLM, maybe.
Asking for a friend.