Is there a way to compile LLVM IR with something other than LLVM? Is this what you mean by output of LLVM front-end output?
I mean that if all top-class talents in compilers technology focuses on llvm there probably wouldn’t be a lot people to be both willing and able to write alternative backends.
For the C++ world the competition by LLVM/clang was fruitful and triggered lots of improvements in gcc/g++. Produced code got faster, diagnostics better etc.
Sure, Fortran is a different area, with less commercial interest a d other challenges. (LLVM is pushed by Apple and Google for non-Fortran needs - it is thinkable that they push decisions, which hinder Fortran, whereas gcc has a different goal and might long term more receptive to Fortran needs?)
I've been thinking a lot about vectorizing loops recently, especially on AVX512 systems. I've mostly been doing microbenchmarks, and I realize that microbenchmarks might not give a realistic full-program view.
LLVM (through Julia and Clang) have a striking performance pattern as sizes vary: https://discourse.julialang.org/t/ann-loopvectorization/3284...
They are fast at multiples of 32, but then performance degrades. This is because it vectorizes loops by creating two loops:
1. 4x unrolled and vectorized (with double precision and AVX512, that translates to 4 * 8 = 32 loop iterations)
2. Scalar loop.
By avoiding that pattern, it was easy to get much better performance at most sizes in a lot of simple cases, like dot products.In wondering about why LLVM's decisions made sense for them, I'm currently leaning towards AVX transition penalties being a big factor. Recently shared on HN: https://travisdowns.github.io/blog/2020/01/17/avxfreq1.html
The thing that struck me is that there is a 9 microsecond period where avx instructions (AVX2 and AVX512) operate at a small fraction of normal speed before the CPU decides to transition to a slower state.
If most of your code is running in L0 (max clock speed) license, then any vectorized code you run into will run at 1/4 speed for about 9 microseconds. If it does run for that amount of time, it'll transition with an 11 microseconds break. It'll have to keep running for a long time to amortize this penalty.
Then, once the function returns to the rest of your scalar code, it'll eventually have to speed up. Basically, large programs are probably fastest if they stay in relatively the same state.
By having a large scalar window, like LLVM does, it's less likely to change. Most loops are probably fairly short, and most code is also scalar, therefore you'll want the CPU to generally stick to scalar mode. Only if loops are very long and likely to take milliseconds would you want them to be vectorized.
Or if they're surrounding by other SIMD code, but that's a sort of global/whole program state you cannot infer while optimizing a single function.
It is likely best to go lean very heavily to one side in your preference of scalar vs vector, but which side is better varies by program. LLVM is essentially leaning heavily toward scalar in their loop behavior, which is probably best for most C and C++ programs. Many Fortran programs might prefer vector. My own (Julia) code does. But I have a hard time talking about Julia or Fortran programs in the abstract. I tend to ensure vectorization. Most programmers don't, so even in these languages, they're likely to prefer and benefit from different defaults than I do.
Actually, it's quite the opposite. You're never going to make money selling a C/C++ compiler, but you can make loads of cash selling a Fortran compiler. It's just that the Fortran compiler is likely to come with your supercomputer.
There is plenty of money to be done in C and C++ commercial compilers, it just depends on the customer base, and features missing from clang/gcc regarding overall tooling experience.