Flang: The Fortran frontend of LLVM [video]
fosdem.org
fosdem.org
On the other I’m a bit worried that once LLVM backend will take off for real, so much focus will concentrate on it, that other compilers (PGI, gfortran, Intel Fortran) development will finally stall or even get dropped and we will wake up in future where LLVM behemoth swallowed all the competition.
This concern is not only limited to Fortran, but also other languages to be clear. Of course I’m very grateful for LLVM project, it’s absolutely great. I think that the idea to focus talent pool from many languages is brilliant and beneficial for everyone. The only problem is when it becomes so focused that it becomes only option. I am very happy for variety and options for compilers and backends we have now.
Another very interesting project, even if it is slightly intimidating in scope, is GraalVM (https://www.graalvm.org).
However, I am excited for Flang and its FIR MLIR-dialect. I haven't benchmarked MLIR-optimized code at all. I'm sure that will change things, but until I test I have no idea by how much.
Intel’s compiler team has actually suggested some patches adding the same mode to LLVM, though I’m not sure what the current status is, since the initial reaction was not overwhelmingly positive.
I think it'd be nice to be able to activate this mode through pragmas.
Does "#pragma omp simd" result in more aggressive use of blocking?
It's also not true that HPC performance is generally dominated by code generation rather than libraries and communication costs, but obviously mileage varies.
I would think margins on CPU sales and on consulting would dwarf that 1% margin, though. Because of that, I think it’s more “winning the benchmarks game, because that’s what sells hardware there” then revenues from compiler sales that keeps intel’s compiler alive.
Is there a way to compile LLVM IR with something other than LLVM? Is this what you mean by output of LLVM front-end output?
I mean that if all top-class talents in compilers technology focuses on llvm there probably wouldn’t be a lot people to be both willing and able to write alternative backends.
For the C++ world the competition by LLVM/clang was fruitful and triggered lots of improvements in gcc/g++. Produced code got faster, diagnostics better etc.
Sure, Fortran is a different area, with less commercial interest a d other challenges. (LLVM is pushed by Apple and Google for non-Fortran needs - it is thinkable that they push decisions, which hinder Fortran, whereas gcc has a different goal and might long term more receptive to Fortran needs?)
Actually, it's quite the opposite. You're never going to make money selling a C/C++ compiler, but you can make loads of cash selling a Fortran compiler. It's just that the Fortran compiler is likely to come with your supercomputer.
There is plenty of money to be done in C and C++ commercial compilers, it just depends on the customer base, and features missing from clang/gcc regarding overall tooling experience.
I've been thinking a lot about vectorizing loops recently, especially on AVX512 systems. I've mostly been doing microbenchmarks, and I realize that microbenchmarks might not give a realistic full-program view.
LLVM (through Julia and Clang) have a striking performance pattern as sizes vary: https://discourse.julialang.org/t/ann-loopvectorization/3284...
They are fast at multiples of 32, but then performance degrades. This is because it vectorizes loops by creating two loops:
1. 4x unrolled and vectorized (with double precision and AVX512, that translates to 4 * 8 = 32 loop iterations)
2. Scalar loop.
By avoiding that pattern, it was easy to get much better performance at most sizes in a lot of simple cases, like dot products.In wondering about why LLVM's decisions made sense for them, I'm currently leaning towards AVX transition penalties being a big factor. Recently shared on HN: https://travisdowns.github.io/blog/2020/01/17/avxfreq1.html
The thing that struck me is that there is a 9 microsecond period where avx instructions (AVX2 and AVX512) operate at a small fraction of normal speed before the CPU decides to transition to a slower state.
If most of your code is running in L0 (max clock speed) license, then any vectorized code you run into will run at 1/4 speed for about 9 microseconds. If it does run for that amount of time, it'll transition with an 11 microseconds break. It'll have to keep running for a long time to amortize this penalty.
Then, once the function returns to the rest of your scalar code, it'll eventually have to speed up. Basically, large programs are probably fastest if they stay in relatively the same state.
By having a large scalar window, like LLVM does, it's less likely to change. Most loops are probably fairly short, and most code is also scalar, therefore you'll want the CPU to generally stick to scalar mode. Only if loops are very long and likely to take milliseconds would you want them to be vectorized.
Or if they're surrounding by other SIMD code, but that's a sort of global/whole program state you cannot infer while optimizing a single function.
It is likely best to go lean very heavily to one side in your preference of scalar vs vector, but which side is better varies by program. LLVM is essentially leaning heavily toward scalar in their loop behavior, which is probably best for most C and C++ programs. Many Fortran programs might prefer vector. My own (Julia) code does. But I have a hard time talking about Julia or Fortran programs in the abstract. I tend to ensure vectorization. Most programmers don't, so even in these languages, they're likely to prefer and benefit from different defaults than I do.
It is likely it will take years if not decades for the llvm based fortran compiler to match the optimization ability of compilers on hpc systems. I also will clarify that the previous statement was an understatement. gfortran is one thing, PGI or Intel compilers won't stall or fall out of use.
Gcc (well g++) is still my primary compiler though I do run my code through clang as well as llvm do catch different bugs. I can imagine llvm becoming my primary compiler at some point, but it will still be a few years.
The evolution of BSD's adoption by commercial entities shows the way.
The more serious problem is the weaponization of open source by the big actors, using it to simultaneously generate a scorched-earth moat around their respective castles while hiding proprietary extensions behind a network connection. I do not consider this to have been good for the software world at all.
As I'm sure you know your prediction has been made for 30 years. Doesn't mean it can't come true but unlike you I don't see the tide moving that way.
The Linux kernel is the only surviving piece of GPL code on modern Android.
Just wait when Fuchsia becomes mature enough.
I mean, it has been around much less time but does not seem to be an especially common choice.
I wrote the library license back in the early 90s because of a similar shift (Unix and Windows were late adopters of the the philosophy of libraries, not just programs, for non- system code, but once they finally started to get on board the GPL had to catch up)
I wouldn't miss the grief associated with ifort, which doesn't live up to the mythology. The salient feature of the PGI compiler is probably the offloading, and I assume that's Nvidia-specific, but I don't know how it compares with current, and upcoming, GCC support. The research computing world would be a better place if the money spent on proprietary compilers sponsored GCC improvements instead.
A further point to make is that LLVM IR is strongly biased towards C, and this is especially true when it comes to the memory model. All memory has to be lowered to access via (essentially typeless) pointers, with optional aliasing qualifiers provided via a noalias parameter attribute (which breaks a lot, because restrict isn't all that common in C), and TBAA. And all higher-order information has to be reverse engineered from this starting point.
C/C++ does not have alias-annotations built-in. Although with template constructs like done in e.g. the Eigen library, alias-optimal code can be generated.
Restrict?
I maintain from long research computing experience (at least back to the days of Alliant) that the rules are highly error prone in practice for users, who frequently deny they even exist, and blame the compiler bugs. (I'm surprised if that's not the case more generally.) I'm not saying they shouldn't exist, or that code needs to contravene them.
What I mean by "default" is that if I declare a function with two array arguments in Fortran, with the simplest possible syntax, the compiler assumes that they do not alias (are not associated). By contrast, if I declare a function with two pointer arguments in C, with the simplest possible syntax, the compiler assumes that they may alias (are associated in some unspecified manner).
The C semantics are certainly safer, but they lead to lots of "Fortran is faster than C" blog posts by people who either don't know about or simply don't want to use the annotations.
That’s when I used it last time :)
Actual standardization (as in, there is a document describing the what's supposed to happen) is helped by multiple implementations because their conflicts will help discover unclear parts. This in return helps new implementations get of the ground if there's a need to produce them. Some standards groups require multiple independent implementation of something to exist before it is allowed to be released as a standard.
Different implementations might be different enough that some changes are easier to make in one than the other. This makes it easier to test these changes, the other implementation(s) can then decide if the gains are worth their effort.
It provides an out if one project resists changes for human/political reasons.
I wouldn't be surprised if there's a few unusual architectures around that use Fortran but aren't handled by LLVM.
That's workable, eg for Python (and in practice, Haskell). But is seen as less than ideal.
The rest of your comment is... way too political towards gpl.
Apple pushed clang as a competitor to gcc much later, closer to 2010-ish.
https://gcc.gnu.org/ml/gcc/2005-11/msg00888.html
Doesn't seem to add up with your proposed timeline. Care to explain?
It was a MODERN Fortran like language that Guy Steele and all were developing at Sun https://en.wikipedia.org/wiki/Fortress_(programming_language...
It did not stand much of a chance once Oracle took over.
Which is a pity, because it had some very interesting ideas, imvho.
(Complaining about downvoting is not typically an acceptable thing on HN but I think this case is meta-respectable, though I appreciate the irony).
llvm: https://go.googlesource.com/gollvm/
gcc: https://github.com/golang/gofrontend
Go: https://github.com/golang/go/tree/master/src/cmd/compile
And I believe all three are maintained by the go team. It should be noted that the LLVM implementation is used by tinygo: https://tinygo.org/
Phoronix article: https://www.phoronix.com/scan.php?page=news_item&px=FC-LLVM-...
I'm new to the Fortran scene, btw. I probably can't do more than fix typos, but the Good First Issues on the github are:
https://github.com/flang-compiler/f18/labels/good%20first%20...
A lot of new interesting languages (Julia, Rust, etc) were empowered by being able to take advantage of LLVM as a backend, so I can't imagine what new languages might spring up around the heterogenous capabilities of MLIR, not to mention existing languages targeting it as well.
The convenience is not really matched by C or C++, where similar features have been added much later by language extensions or 3rd party libraries, resulting to more complicated usage, fragmentation, and interoperability problems. Newer languages also have similar issues, so for the user base that uses Fortran, there's a lack of viable competitors.
If you want to do common numeric operations they might be a FORTRAN code that is battle tested and performance tuned and it usually not hard to call from C, Java, Python or some other language.