Intel's “cripple AMD” function (2019)
agner.org
agner.org
A serious question - how many common benchmark packages are compiled by ICC or uses Intel MKL? I hope the number is limited, otherwise all the benchmarks published by mainstream PC reviewers are potentially biased. If there's a serious ICC-in-benchmark problem, then only Phoronix's Linux benchmarks are trustworthy - the majority of benchmarks on Phoronix uses free and open source compilers and testsuits, with known versions, build parameters and optimization levels. Thanks Michael Larabel for his service for the community.
https://www.archer2.ac.uk/ will run a lot of that sort of thing. I think at least cp2k and CASTEP are included in the benchmarks, but they're listed somewhere.
> I have compared other DFT code on 64-core Bulldozer with 12-core Sandybridge nodes
And what's that comparison supposed to tell us, aside from the obvious fact that MPI introduces latency? That's just about the number of cores, not the performance of a each core. You need to compare 64-core AMD node against a 64-core Intel node.
For a recent exercise, spending rather more money on AMD CPUs for the UK Tier 1 system, look at the Archer2 reference and benchmarking for it. It's expected to run large amounts of VASP-like code; www.archer.ac.uk publishes usage of the current system. Circumstances differ, and I'm pointing out contrary experience, understanding the measurements and what determined them.
[1] https://numpy.org/doc/stable/reference/routines.linalg.html?...
Small matrix multiply benchmarks on a Zen2 (Ryzen 7 4700U), featuring MKL 2020.1.216+0, OpenBLAS, and Eigen: https://gist.github.com/stillyslalom/bd916e3d26b4531364676ac...
MKL's benchmark performance requires AVX + FMA. 3.5 GHz * (4 add + 4 multiply) * 2 fma/cycle = 56 peak GFLOPS. To exceed 50 GFLOPS without them would imply the CPU ran at 12.5 GHz.
OpenBLAS, on the other hand, actually performed poorly because it was limited to SSE thanks to a bug preventing the CPU from being recognized as Zen2.
Odd. I am trying on my 3700X and it is definitely not using AVX, FMA or AVX2 code paths. Intel MKL 2020 update 2:
ldd ~/git/sticker2/target/release/sticker2 | grep mkl_intel
libmkl_intel_lp64.so => /nix/store/jpjwkkv1dqk4nn8swjzr5qqzp0dpzk2f-mkl-2020.2.254/lib/libmkl_intel_lp64.so (0x00007fe786862000)
I checked the instructions in with perf and it is using an SSE code path. Also, as reported elsewhere, MKL_DEBUG_CPU_TYPE=5 does not enable AVX2 support as it used to do. julia> using LinearAlgebra
julia> BLAS.vendor()
:openblas64
julia> BLAS.set_num_threads(1)
julia> peakflops()
3.9023447970402664e10
julia> using LinearAlgebra
julia> BLAS.vendor()
:mkl
julia> BLAS.set_num_threads(1)
julia> peakflops()
4.8113846984735275e10
That's close to the ~50 Gflops I saw in @celrod's benchmarks.I have now also compiled the ACE DGEMM benchmark and linked against MKL iomp:
$ ./mt-dgemm 1000 | grep GFLOP
GFLOP/s rate: 69.124168 GF/s
Most-used function is mt-dgemm libmkl_def.so [.] mkl_blas_def_dgemm_kernel_zen
So, it is clearly using a GEMM kernel. Now I wonder what is different between PyTorch and this simple benchmark, causing PyTorch to result in a slow SSE code path.Conclusion: MKL detects Zen now, but currently only implements a Zen code path for dgemm and not for sgemm. To get good performance for sgemm, you have to fake being an Intel CPU.
Edit, longer description: https://github.com/pytorch/builder/issues/504
FWIW, on my [Skylake/Cascadelake]-X Intel systems, Intel's compilers performed well, almost always outperforming GCC and Clang. But on Zen, their performance was terrible. So I was happy to see that MKL, unlike the compilers, did not appear to gimp AMD.
It's disappointing that MKL doesn't use optimized code paths on the 3700X.
I messaged the person who actually ran the benchmarks and owns the laptop, asking them to chime in with more information. I'm just the person who wrote that benchmark suite.
When OpenBLAS identifies the arch, it is competitive with MKL in single threaded performance, at least for matrices with a couple hundred rows and columns or more. But MKL truly shines with multiple threads, so scaling on a 32 core system would be interesting to look at.
Explained in the NumPy docs here: https://numpy.org/install/
This is definitely understating MKLs market share in BLAS. It's an extremely common BLAS backing for Python libraries.
In addition, software like Mathematica/Matlab which use only MKL can affect decisions even for office workstations.
When run on an AMD, any program built with Intel's compiler should have the environment variable set. I don't think there is any downside to leaving it on all the time, unless you are measuring how badly Intel has tried to cripple your AMD performance.
https://en.wikipedia.org/wiki/Math_Kernel_Library#Performanc...
This is quite bad, since a lot of software relies on Intel MKL as the default BLAS implementation (e.g. PyTorch binaries).
It may even be easier to replace the function altogether with LD_PRELOAD.
Make a file with the following content:
int mkl_serv_intel_cpu_true() {
return 1;
}
Compile gcc -shared -o libfake.so fake.c
Run LD_PRELOAD=libfake.so yourprogram
And it uses the optimized AVX codepaths.Disclaimer: may not be legal in your country. I take no responsibility.
patchelf --add-needed libfakeintel.so yourbinaryOf course, the second one makes no sense as x86 programs run just as fine on AMD as Intel with the same feature set (albeit at different speeds)
Is there something I can add to my bashrc to handle that?
If that fails, as OP implies, you can still override the function by creating a tiny library with it always returning true. On GNU/Linux systems, you do that using LD_PRELOAD. Perhaps someone's already done that so you just need to download, compile and set it.
Sorry for the lack of specifics, but I do not deal with these libraries, yet I was still hoping to point you in the right direction.
> Avoid the Intel compiler. There are other compilers with similar or better performance.
This is not really true IMO, but even as an aside, the Intel compiler has the enormous advantage of being available cross platform. So we can use it on Linux and Windows, and provides MPI cross platform. We upgrade fairly regularly and that provides us with less work.
My own tests found that PGI compiler performance was worse than Intel for C++, and that now appears to have been discontinued on Windows anyway with NVidia's new HPC compiler suite replacing it. GNU can run everywhere, but performance is around 2.5x worse on Linux for our application use case because it doesn't perform many of the optimisations that Intel does. We use MSVC on Windows just because everyone can have a license, and performance is much worse.
The other thing is that MKL is pretty stable and gets updated. If I use an open source BLAS/LAPACK implementation - sure, it works, and it may even give better performance! But it's not guaranteed to get updates beyond a couple of years, and plenty of implementations are also only partial. We pay Intel a lot of money for the lack of hassle, basically.
I don't know about MKL stability, but reliability definitely isn't something I associate with the Intel Fortran compiler (or MPI) in research computing support.
How does this advantage not apply to gcc? Isn't gcc the most cross-platform compiler ever?
At which point, its not just the compiler (which GCC is pretty good at), but also the threading implementation (which I can believe that GCC has an inferior Windows-threading OpenMP implementation).
I don't really use either tool. But OpenMP + GCC on Windows doesn't sound like it'd be fast to me.
--------
MSVC only has OpenMP 2.0 support (OpenMP is all the way up to 5.0 now).
OpenMP, despite being a common interface, also is pretty reliant on many implementation details for performance. One way of doing things on GCC could be faster than another, while it could be the opposite on ICC. Its quite possible that their codebase is tailored for ICC, and that recompiling it under GCC (with a different OpenMP implementation) results in weaker performance.
I wouldn't expect 250% performance difference in normal code however. GCC and ICC aren't that far off under typical circumstances.
Free BLASs are pretty much on a par with MKL, at least for large dimension level 3 in BLIS's case, even on Haswell. For small matrices MKL only became fast after libxsmm showed the way. (I don't know about libxsmm on current AMD hardware, but it's free software you can work on if necessary, like AMD have done with BLIS.) OpenBLAS and BLIS are infinitely better performing than MKL in general because they can run on all CPU architectures (and BLIS's plain C gets about 75% of the hand-written DGEMM kernel's performance).
The differences between the implementations are comparable with the noise in typical HPC jobs, even if performance was entirely dominated by, say, DGEMM (and getting close to peak floating point intensity is atypical). On the other hand, you can see a factor of several difference in MPI performance in some cases.
gcc has always produced faster code for at least 15 years. In fact, it is the Intel compiler which has caught up in the most recent version.
By nature, this code had no usage of floating point in its critical path.
I haven't bothered with icc in years though.
One thing I found that you do have to be careful with though is ensuring that Intel uses IEEE floating point precision, because by default it's less accurate than GCC. This causes issues in Eigen sometimes, we ran into an issue recently after upgrading compiler where suddenly the results changed and it was because someone had forgotten to set 'fp-model' to 'strict'
It goes without saying that you should use -O3 (or -O2 for some rare cases) otherwise. I am mentioning it just in case because 2.5x slower sounds so exotic to me that the first intuition is that you're omitting important optimization flags when using GCC. GCC was faster than Intel on everything I tried in the past.
I don't know if Oracle is still using ICC for that or not. (If you download Oracle RDBMS, and check the binaries, you will be able to work it out. I can't be bothered.)
[1] https://www.businesswire.com/news/home/20030507005238/en/Ora...
Many compilers statically link implementations of various built-in functions into the resulting executable, and that can result in different symbol table entries
I can't even find benchmarks of ICC vs a current GCC but they were pretty even the best part of a decade ago. GCC is a mess compared to LLVM but it's quick.
I'd suspect the same from Unity.
Now that AMD EPYC processors are powering a lot of next generation super-computer clusters, we're going to have to figure out some workarounds!
to here: https://github.com/oneapi-src/oneDNN/blob/master/src/cpu/x64...
It's using feature-flag checks, not family checks, so you shouldn't be affected if you're using oneDNN.
Anecdotally, (ignoring that I'm still not sure whether to trust it or not) Intel stopped developing IACA and suggested (but not recommended) LLVM's MCA - which does suggest a changing of the guard in some way.
Edit: the link I posted follows Agner's advice from the bottom of OPs link. However I think the extra information that it adds is that Zen2 Threadrippers outpaced then-current Intel's top contender. Once Zen3 and Intel's 11th Gen become available, repeating this benchmarks would be very valuable.
BLIS is used by AMD and is a good open alternative to MKL (for BLAS) across many platforms. https://github.com/flame/blis/blob/master/docs/Performance.m...
I wonder if this Intel compiler 'superiority' is still the case today, or if this is just a meme at this point.
For C based software, its a much closer run thing, and often sticking with GCC avoids weird segfaults when mixing Intel and GCC-compiled Linux libraries.
To be generous...
Where do you typically see lack of inlining and vectorization with GCC? I'm curious because most times people have said GCC wouldn't vectorize code that I've been able to try, it would, at least if allowed -ffast-math a la Intel (as in BLIS now).
[1]https://developer.amd.com/amd-aocl/blas-library/
[2] https://developer.amd.com/open-source-strikes-again-accelera...
It is a myth that icc produces faster binaries that may have been true 25 years ago.
You can do the same when your machine has non-intel CPUs that are supported by a lot of compilers. If you are on power9 or arm the compiler list gets shorter. And a lot of supercomputers start to contain accelerators (often, but not always Nvidia GPUs) in which case there is often only one supported compiler and you have to rewrite code until that compiler is happy and produces fast code.
"The same effect was documented with some of the most popular mathematical software packages, including Mathematica, Mathcad, and Matlab."
It's a bit worse than that. Intel has no obligation to support optimizations that aren't unique to AMD; they're allowed to disable SIMD extensions that AMD processors declare support for, while at the same time using all of those SIMD extensions on Intel CPUs. They just have to include the disclaimer that their compiler and libraries may be doing this.
It’s straight up anti-competitive, but consumers aren’t smart enough to understand that it’s a problem; A consumer just sees biased benchmarks that show Intel outperforming AMD, and then choose Intel.
As far as I understand it quite the opposite, it explicitly mentions that it may not apply "optimizations that are not unique to Intel" to other processors. It wont select the optimal code path unless the CPU vendor ID is set to GenuineIntel and fall back to the worst path your compile settings include.