How Much Does a Compiler Cost?
embecosm.com
embecosm.com
Typically in compiler development, there's massive R&D time, many hours spent analyzing memory & CPU patterns to figure out which code to emit. You also take a look at disassembly of generated code and nowdays perform data-mining to figure out which patterns of optimize for (typically on non-kernel or hotspot-heavy type workloads).
The end result of alot of investigation can be few lines of code which have to be absolutely robust and production ready at the time of release under standard configurations (i.e -O1, -O2 etc...)
Its a costly piece of engineering which is why some companies still charge money for it.
I'd say the article's estimates are accurate for a simple compiler that emits results at say -O0 or -O1 level of efficiency, but anything higher will start to require more and more resources with diminishing gains over time. But this can be said about most things in engineering.
Unless I'm misunderstanding "by default", this isn't true. LLVM is language agnostic and supports many languages - even the primary C-like frontend Clang supports more than just C/C++.
Much like how the older versions of GCC used to just produce ASM files to be ingested into GAS to create the binary objects.
There exists no such thing as "language agnostic". Every intermediate representation contains properties (very often implicitly) that are specific at least to a class of programming languages.
If you don't believe me, here is the code that optimizes based on this.
https://code.woboq.org/llvm/llvm/lib/Transforms/InstCombine/...
https://code.woboq.org/llvm/llvm/lib/Transforms/InstCombine/...
Some highlights:
Now, before GHC generates assembly code (abiding by the calling conventions previously described), it also needs to optimize it. It optimizes a form called "Cmm", which is like a very low level compiler language for performing optimizations. This language does include functions.
When you use the LLVM backend, GHC translates every Cmm function to an LLVM function. In order for everything to work out, we patched LLVM so that functions can have an annotation saying they follow the GHC calling convention.
...
What this means is, we have completely side stepped LLVM's support for GC.
So the support mostly comes from the Haskell side, by compiling to the C-like Cmm (I think that name comes from C--, the idea being a language between assembly and C in abstraction level) and patching or otherwise working around the part where LLVM's feature set wasn't appropriate.
LLVM Intermediate representation in practice is C++, AST is C++. LLVM do not support many languages,many languages support LLVM, which is different.
That is many people, or languages' designers, have made the necessary work to make their languages compatible with LLVM IR. If the language is similar to C++, this is a simple task, if it is very different, it is very hard.
LLVM does not support other languages, in fact, codebase changes a lot making maintenance painful.
Specifically, in my limited experience LLVM IR resembles typical hardware rather than C++. Modeling the hardware is a design goal for C++, of course, so anything hardware-like must be C++-like. Can you describe some ways in which IR resembles C++ that are not plausible ways to resemble hardware? I'm really curious.
We now assume ARM as being a default necessary platform to support, 20 years ago it really wasn’t clear
Making sure everything we wrote was compatible with the public tree was expensive. In the end, one of the last things I did saved a lot of money and improved code quality for everyone: I forked gcc from the FSF (and made the egcs tre the de facto, and eventually de jure version).
Nowadays forking is considered routine, and even beneficial, and projects have steering committees that are responsible for master releases. At that time, forks were considered a tragedy, and it took me over 6 months of discussion with various people to build a reasonable consensus (you can see the people's name at the bottom of the letter I wrote announcing it ( https://gcc.gnu.org/news/announcement.html ). It's been a good model. Note that RMS's name is not on the letter -- he was furious and was sure this would destroy the FSF. Instead it has strengthened it.
By the way, to compare these numbers: Red Hat went public with, AFAIK less than 10 engineers and when Cygnus and Red Hat merged, though both companies had about 200 people, Cygnus had over 160 engineers while Red hat had a couple of dozen. Of course we had a lot more revenue, and in the years since RHAT has invested heavily in free software development. But their company value at that time was something like 3X ours -- a very valuable lesson.
In 5 years we'll all be using risc-v with a nice DSP extension anyway... And after typing that I felt the need to google: https://riscv.org/wp-content/uploads/2016/07/Wed1000_dsp_isa...
[1] At least it was 18 months ago when I still working in the space, can't imagine that's changed since.
Also, it’s been a few years since I touched IBM’s POWER hardware with AIX, but I’m pretty sure XLC and whatever the Fortran compilers were didn’t come free.
I think I also recall something about gfortran deliberately choosing some more conservative optimization strategies, which could also explain some differences. Or I could be misremembering absolutely everything :-)
GCC does correct optimization by default (modulo bugs). ifort's default set of flags on GNU/Linux at least used to be non-standard-conforming (maybe fast-math?). Comparisons need to use equivalent flags for gfortran. I don't have the ones I once noted to hand, but they'd include -Ofast -funroll-loops -march=native. The ifort default also generates static binaries, which require static libraries and which you typically don't want so that you can do profiling, for instance.
And to be clear, I’d love to see Free and Open Source options be the standard for HPC. It sounds like we may be much closer to that state today than we were years ago when last I worked with people highly concerned about performance in this area. That’s great news.
There used to be a published set of results from Polyhedron -- presumably to help to sell proprietary compilers -- which had a geometric mean favouring ifort by ~20%, if I remember right. However, the last I saw didn't use the latest GCC, required the same flags for each case (which didn't include profile-directed, for instance), and the cases with the biggest differences actually bottlenecked on libraries, not generated code. gfortran is infinitely faster on most architectures, of course.
Most benchmark results I see aren't useful because they don't supply the parameters needed to be reproducible and they don't have profile information to allow you to understand, and maybe improve, them.
If you think this cheesy, consider this: you are a manager of Intel compiler team. Someone implemented an optimization that improves performance by 1% on Intel, but 10% on AMD. Do you think it is a good change? The change will be no-brainer on GCC, but consideration for ICC is necessarily different.
Now I want to write compilers and developer tools, but there's no money to be made.
It's sad. One the one hand it's great that we have access to so many amazing compilers for free. But now there's a monopoly of compilers - and by extension, programming languages - by large corporations. They're all kind of the same thing too - vaguely C-like statically typed ALGOL languages with OO and FP features bolted on top. Swift/Kotlin/C#/Java/Typescript/Go, it's all much the same.
I don't think anything novel (a smalltalk, prolog, haskell or scheme) could get any mindshare today, unless a big company lost their mind and tried to push something weird.
I think there's money to be made, but there's a lot of competition and you need a very specific niche to break through.
Nowadays my company pays volume licenses for MSDN.
Also, I would argue that when one buys an Apple computer, given the price difference to PCs, the price of the development tools is included.
I also remember dad having to drive me to a university in a bigger city so that I could buy VC++ from the university there.
I wish I knew about linux back then.
Well, that's certainly the ideal case. I can only wonder what historic defect rates have been for commercial compilers.
A great benefit of GCC was always enabling learning and research, though there are more opportunities these days.
For example, how much is rust? python?
It's weird how a compiler can cost so much, yet it can still produce unsafe code because the language is not safe by design. In that optic, rust might save a lot of money.
My strategy was to start with a very small grammar and grow it.