Hurrah for competition!
(of course, neither GCC or clang could _ever_ beat ICC at some of my tests.. grumble)
They never tried, honestly - ICC had plenty of defaults that were targeted at performance above correctness.
That's fine - but it wasn't what either clang or gcc was going to go for. Even when trying to compare apples to apples, folks often compared them by trying to get GCC/Clang to emulate the correctness level of ICC (IE give GCC/Clang more freedom), rather than the other way around :)
Which, again, similarly understandable, but also a thing that you'd have to spend a while on in GCC/Clang to get to a reasonable place)
Small-number-of-target compilers like ICC are also fundamentally easier. Lots of techniques (IE optimal register allocation[1], etc) that add up to performance gains are a lot more tractable to really good applied engineering when they don't have to be so general.
Similarly small-target-market compilers are also easier. Over time, high performance was literally ICC's only remaining market. So they can spend their days working on that. GCC/Clang had to care about a lot more.
[1] This happens to now be a bad example because advances finally made this particular thing tractable. But it took years more, and feel free to replace it with something like "optimal integrated register allocation and scheduling" or whatever is still intractable to generalize.
> They never tried, honestly - ICC had plenty of defaults that were targeted at performance above correctness.
Is it possible to get ICC level performance out of open source tools? Much of my CPU bound work relates to array signal processing, which, if you can code it right :) lends itself heavily to SIMD branchless pipelines. Plus some scatter gathers on a group of other cores to calculate a sparse crosscorrelation tensor.
I would love to be able to get same or better performing code out of clang or gcc, even it it takes 2-4x more work than with ICC...
Or at least, we extensively tested this at Google, before, during, and after the move to LLVM.
On many thousands of libraries, binaries, etc, made up of hundreds of millions of lines of C++.
While there were wins and losses, on average, LLVM was a consistent net positive.
That was true despite having spent 5+ years of having a large team dedicated to doing nothing but finding places to improve performance of GCC compiled code (and contributing back patches), and doing so very successfully.
That said, as time approaches infinity, the compilers are going to generate the best code for the things that someone took the time to analyze and make the compiler better at.
There is, in the end, no magic that makes GCC better than LLVM or vice versa. The vast majority of it is tuning and improving things little by little for whatever targets someone is trying to improve.