Also: if icc is going to go with LLVM as a backend, then what is the point of using icc at all? Why not just use clang?
Also: if icc is going to go with LLVM as a backend, then what is the point of using icc at all? Why not just use clang?
From the blog post:
"Not all our optimization techniques get upstreamed—sometimes because they are too new, sometimes because they are very specific for Intel architecture. This is to be expected and is consistent with other compilers that have adopted LLVM."
I don't see any good technical reason in this marketing language though, tbh:
- "Our optimizations are too new" -> So put them behind a feature flag when upstreaming? And why inflict them on icx customers if they are "too new"? What does that even mean?
- "Our optimizations are architecture specific" -> ...yes? Emitting good architecture-specific code is the whole point of an optimizing compiler? How is that an argument against upstreaming?
- "Other compilers also don't upstream everything" -> That's a non-argument.
I get it, there's no money to be made in just implementing all your optimizations in clang directly. But these reasons (together with the repeated emphasis of how they do contribute to LLVM) seem silly. To be clear, I'm all for Intel helping improve LLVM, but the technical arguments for maintaining a separate commercial version don't convince me at all.
The API and fundamental part of LLVM is its IR, LLVM-IR, which is where most optimizations happen.
From this point-of-view, LLVM is a "platform", and an extremely brittle one: every release has breaking changes to the IR, the IR is constantly evolved to support new hardware and new optimizations, etc.
When you build a tool on top of LLVM, you are buying into this rapidly changing platform.
The only proven way of using the LLVM platform competitively is to be part of its future: follow upstream closely, upstream most of your code, and actively participate in the platform evolution so that your competitors can't turn it against you.
If you keep most of your code private, following upstream gets very hard. You skip one release, and then its 10x harder, so you skip another release and stop participating in LLVM's evolution cause you'll have to wait years for changes to upstream to land on your compiler. Your competitors do what's best for them, and those can be things that are bad for you, and then you are proper screwed, cause you can't migrate away from LLVM either.
Companies like Intel and NVIDIA do this, e.g., nvcc and ISPC are stuck on LLVM 7 (~4 years old), but these companies have built huge technology stacks like CUDA or DPC++ on top of it!
Intel and NVIDIA might have enough manpower to maintain their own outdated fork of LLVM forever, but at some point it just stops being LLVM, and these companies are not really much better off than where they started.
IMO, building all your technology on top of a platform that either is or can be under your competitors control is just a really bad idea.
One would hope that these companies would realize this and contribute back and help the community as much as possible, but in practice they just don't. By the time they realize it, it's already too late.
The joy of permissive licenses!
Android is a great example. Typical Qualcomm/Mediatek/etc. behavior: take a "stable" kernel, stuff it with custom undocumented drivers and junk, be stuck on that version forever. The only thing the GPL changes is the vendor posts a source dump of dubious usefulness in some obscure section of their website.
I don't think that's such a great example in this context.
Qualcomm, Mediatek, etc. are hardware vendors after all. Software to them is a necessary evil, not a reason d'être or complementary tool (as it is for NVIDIA and Intel). Their customers fall in the same category - smartphone manufacturers want to sell units, not keep software up-to-date.
https://www.msn.com/en-us/money/technologyinvesting/nvidia-i...
It's a very similar story for Intel. NVIDIA and Intel are selling hardware because of their software portfolio. How many GPUs would NVIDIA sell in the HPC market if it wasn't for CUDA and the various support libraries around it?
Intel's non-client solution revenue was over 40% of their total revenue in 2020. This includes their (software-) AI solutions, applications, licence business and services. So Intel, too makes a significant amount of money from software and services around their software ecosystem.
We don't have to hypothesize about this. The HPC GPU market has multiple vendors, just look at how many HPC GPUs AMD or Intel are selling: AMD and Intel have ~0.8% or so market share. NVIDIA has >99%.
Pretty much every review of HPC GPUs states that AMD GPUs are both faster and cheaper.
So how come they don't sell?
The answer is software.
That said, this "junk" dump can still be used by the community of users to upgrade themselves the software on that hardware, so that's still a nice improvement.
Back to LLVM the question is whether these companies decide to not contribute upstream because they don't bother to make clean patches or because they want to keep it for themselves.
I'd argue that they would be much better off long-term making clean patches anyway, so that's not a valid reason for not contributing.
And even if theynjust dumped their patches, the community could still take them and incorporate nice optimization into upstream.
>If you keep most of your code private, following upstream gets very hard. You skip one release, and then its 10x harder, so you skip another release and stop participating in LLVM's evolution cause you'll have to wait years for changes to upstream to land on your compiler. Your competitors do what's best for them, and those can be things that are bad for you, and then you are proper screwed, cause you can't migrate away from LLVM either.
This is an interesting take. I've heard the claim that GCC kept its codebase cryptic to prevent companies running home with it and not upstreaming changes, maybe that's LLVM's strategy.
It turns out that it is impossible to convince people working on LLVM on their free time or academics to invest part of their time on "preserving" API compatibility instead of adding new features, fixing bugs, or improving perf, for the benefit of companies that don't want to contribute their improvements to the community.
https://github.com/freebsd/freebsd-ports/blob/85cccf4f15c42d...
But hopefully that will be fixed soon.
Intel's C compiler has always played games with non-intel x86 architectures compared to their own.
Sounds really shortsighted. I would imagine revenue from Intel dev tools is a rounding error in Intel's complete revenue stream?
ICC Classic is legacy/dead and also free now. The LLVM-based DPC++ compilers are free as well. These are only making money indirectly and through support contracts.
There are pros and cons for sure. I'm not sure what it best.
Based on (minor) personal experience of gcc forks, it's not unusual to make some fairly significant change to some major data structure to make your CPU work better, but which would break several other backends.
There may not be any nice way of integrating these changes in a way which would make the acceptable upstream, without multi-months of work refactoring huge chunks of the compiler (which would still also need lots of work on all those other architectures, to make them compatible with your changes).
Like the famous check-for-intel-model-instead-of-feature-flag optimization? [0]
This is exactly why GCC is GPL and why RMS didn't want to make it more modular. Taken to the extreme we could end up with proprietary hardware that requires a proprietary (closed) compiler even though it's built on open source. Going back to "trusting trust" things might not be so good, and we know Intel is happy to build untrustworthy chips.
However, that is an actual argument. This isn't GCC where that would be required. This is LLVM where that is allowed. This is a choice, Intel's choice, because LLVM allows that, and it's probably at the corner case level of an Intel-specific icc-specific feature, as in, not particularly generally useful.
There is a benefit of this for normal LLVM development as well. It means that Intel is responsible for maintaining it and normal LLVM developers aren't. If I do something that breaks something in the LLVM AMD GPU backend, that's on me. If I do something that breaks something in Intel's code, that's on them.
Now that they're using LLVM they should just upstream everything to the LLVM project and quit selling icc.
They kind of already did. And it's on LLVM, too.
It is not like clang is enjoying C++ Builder RAD abilities, PS 4 and 5 optimizations, or bitcode format used by watchOS.
Companies can keep their optimizations in secrets. Just like why Playststion's OS is based on BSD.
Looking at output assembly language won't tell me anything about the algorithms behind solving the register allocation problem (aka: live data problem, which is a knapsack problem and/or graph-coloring problem IIRC).
I'd be able to see that yes, compilers are good at deciding which registers should hold which data. But that's not sufficient at actually learning how the algorithm / register selection process works.
I assumed the icc secret sauce was in the backend: deep knowledge of how each and every single operation is implemented in every uarch.
Maybe they've written their own x86 backend for LLVM, and are using the rest of LLVM for IR->IR transformations and vectorization?
Although the diversity of C/C++ compilers is descending, the number of compiler backends in general is still pretty high (Go, .NET compilers & runtimes, Java compilers & runtimes, JS engines, etc), and LLVM can’t fill all the niches. I don’t think that we’re at a point where research is stifled.
Browsers are the most oft cited example of monoculture issues be it the old days with IE or the new days with Chrome, I'm surprised you haven't run across comments on that one over the years.
They're not upstreaming all their optimisations. I'd be surprised if they upstreamed all their FPGA support as well.
That's quite reassuring for users of LLVM on non-Intel hardware, where ICC "optimisations" turn into pessimisations.
In it, the guest talks about how he couldn't rely on LLVM for things like optimizing array operations; LLVM apparently does a poor job of that so he had to implement his own. Given that one of the key selling points of Intel's compiler is that it does a better job with SIMD optimizations, it may be exactly the same story here.
The LLVM IR doesn't understand array operations, so the Fortran frontend must 'scalarize' the array expressions, that is, turn them into the equivalent loops. Only after that can the LLVM middle and back-ends try to vectorize that scalar IR for execution on SIMD (or vector, or GPU) HW. The problem is that this scalarization loses some information, and thus for good performance on array code the Fortran frontend must implement some optimizations on array operations before scalarizing.
There is an LLVM subproject called MLIR (https://mlir.llvm.org/ ) that aims to build a higher level IR that understands arrays, and can be useful for things like optimizing deep learning graphs, but also things like Fortran frontends could make use of it. AFAIK the flang Fortran project aims to make use of it, but I haven't followed development that closely.
I was also under the impression GCC still outputs better asm than LLVM overall.
Conversely, I recall reading the zig team ran into some woes with the latest LLVM release, and that there's a lot of churn in LLVM APIs from version to version.
(I'm legitimately curious)
My feeling is, as soon as they can be sure they have achieved 100% binary compatibility, MS will jump. The optimizers in VC++ are light years behind, and companies like Apple are embarrassing them on the compatibility front with things like Rosetta 2 (which is a much harder problem to solve)
GCC hasn’t remained stagnant though.
So basically, if Stallman hadn't have missed the message, there's a very real possibility that LLVM would have ended up licensed under GPLv2 & then re-licensed to GPLv3. The engineering world would probably look radically different if that had happened.
Azure Sphere OS is also GCC only, despite Microsoft's new foundled love for clang.
I'm fairly confident the FSF would keep it alive for a long time, but I don't see why it would be necessarily be a priority for the Linux devs to keep GCC forever.
Android Linux kernel fork actually builds with clang, and Linux kernel is yet to accept the changes made by Google.
No Rust in GCC? No play.
Rust is only following down the footsteps of D and Go.
I would rather see both of these compilers to stay competitive to push themselves higher.
Those changes just don't get to upstream.
https://www.kernel.org/doc/html/latest/kbuild/llvm.html
I could be missing something, but I don't see any suggestion that you need a specific forked tree with patches to build with LLVM, and I've seen people filing bugs about using LLVM sanitizers to build the vanilla tree, so I don't think the expectation is that you need to apply a huge out of tree patchset for this to work any more?
My latest update was that not all patches were accepted upstream, or Google didn't care about upstreaming them, whatever.
There are some Linux Plumbers talks, or from Linaro, about this a years back.
AFAICT from the issues page, Clang and binutils/LLVM tools work fine with no patches for the mainstream archs and when not trying to be super-fancy with custom flags. The more non-mainstream one goes with arch or flags the more likely one will run into something.
[0] https://github.com/ClangBuiltLinux/linux/issues (Note they use github for issues/wiki, not code, so no surprise the 'linux version' in code is oldish).
But I see no signs of MSVC going away any time soon; Microsoft has a very active and capable compiler team.
Because clang is painful to use.
The entire point of GCC and the GNU project is to have an entire system with only free software, and the way to achieve that is not by allowing in more proprietary code.
Let GCC die if it has to die, but changing the license would make it meaningless.