Intel MKL on AMD Zen
danieldk.eu
danieldk.eu
* Intel's “cripple AMD” function (2019)"
https://software.intel.com/content/www/us/en/develop/tools/m...
Also, I would like to push for an anti-acronym day. A full day where you have to use full and complete terminology. Ironic thing is, you don't add that many syllables when you say "Proof of Concept" vs POC and similar such absurd overused abbreviations.
Use of acronyms is usually an audience dependent thing. I am frequently in the in the audience of talks/papers where I don't know all the acronyms. I think the general rule should be, for writing, if there is a reasonable percentage (say more than 10%) of readers who don't know the acronyms, define it on first use. When speaking, give some context or define the terms early in your talk, to help folks along.
I say in a meeting the other day with marketing folks. They may be the worst offenders. Techies are pretty bad too.
I don't agree. The article quite clearly refers to Intel MKL, which is widely known at least in the number crunching world. Plus, it's just a blog post.
https://www.reddit.com/r/Amd/comments/ik4bt9/intel_recently_...
Also, does anyone know if one can use patchelf on e.g. Python (or the numpy compiled section?) to get MKL/Zen support? I don't have my Zen CPU on me to test.
1. They are sharing increasingly much code between oneDNN (formerly MKL-DNN) and Intel MKL. So, it simply pays of from a development cost perspective.
2. Intel sees AArch64 as a serious competitor. AMD may be an enemy, but at least an enemy on the same architecture. Better to strengthen x86_64 than to give AArch64 too much momentum in HPC. (The #1 supercomputer in TOP500 is AArch64 [1])
3. Write Zen kernels, make them slightly less efficient than the Intel ones. They still beat OpenBLAS et al. in benchmarks, even on Zen. But to HPC customers they can show that Intel CPUs are faster.
I think (1) is somewhat optimistic, given the huge benefit that they had with MKL in HPC. If (1) is not the case, I definitely hope (2) is.
[1] https://www.top500.org/news/japan-captures-top500-crown-arm-...
Maybe?
> Maybe?
I guess it depends on where you work. Most places would probably fire someone who "took the initiative" to improve a competitor's product on company time.
[0] https://pharr.org/matt/blog/2018/04/29/ispc-retrospective.ht...
[1] https://devmesh.intel.com/projects/intel-ispc-in-unreal-engi...
Historically, this makes the most sense to me: AMD has pooped the bed before, and Intel was there to pick up the pieces and reclaim its crown. I think there's a legitimate worry that if ARM takes enough marketshare from x86/x64, that Intel will never get it back.
> 3. Write Zen kernels, make them slightly less efficient than the Intel ones. They still beat OpenBLAS et al. in benchmarks, even on Zen. But to HPC customers they can show that Intel CPUs are faster.
IMO, this is more on AMD than Intel. If Intel's kernel is faster than the open alternative, then Intel is doing AMD a favor, even if they're not doing the absolutely best job they possibly could.
If AMD chips are slow on the Intel-optimized kernel, then it's partly on AMD.
If Intel writes an "AMD-optimized" kernel that's worse on AMD than the "Intel-optimized" kernel, that's definitely not on AMD. And in that case it's not doing them a favor to make that kernel when they could just use the same code on everything.
" ... Intel must not ... Intentionally include design/engineering elements in its products that artificially impair the performance of any AMD microprocessor."
What we are seeing has to be a new deal.
This applies both to CPUs (where they have to rely on Intel for this) and toe GPUs (where their offerings are under resourced, buggy and lagging what NVidia offer).
I'm a true believer in ROCm HIP. It has tremendous potential. I joined the company specifically because I wanted to help ensure its success... mostly by writing fast, reliable code, but also by listening to our users.
If you (or anyone else) would prefer to respond privately, my email is my HN username at gmail.
how is that good news when your investigation shows in reality its "Intel seems to be adding cripple-Zen kernels" when compared to spoofing Intel? 382 GF/s vs 430 GF/s
The entire Industry should develop a true interest to push OSS alternatives. And all OSS projects should honestly think twice about implementing such a CSS like the MKL as a standard.
Maybe it is faster but less accurate to run the intel kernel on AMD hardware.
There are results for AVX2 systems with older versions of OpenBLAS and BLIS, which is AMD's BLAS at https://github.com/flame/blis/blob/master/docs/Performance.m... I haven't seen or made measurements for EPYC2 yet (and don't have results with the current versions online for Haswell and SKX) but I'd be surprised if AMD BLIS doesn't perform equally well on EPYC2. OpenBLAS serial DGEMM is very similar to MKL on SKX, for instance, with BLIS not far behind. I think AMD work has contributed to BLIS Haswell performance; their BLAS certainly supports Intel hardware decently, as well as aarch64, at least.
If you're interested in small dimension matrix multiplication on AVX2 hardware, consider libxsmm and AMD's recent "SUP" support in BLIS, e.g. https://github.com/flame/blis/blob/master/docs/PerformanceSm... MKL only got good at small dimensions because of libxsmm.
Off-topic for AMD, but in lieu of detailed figures, here are the first few points for measurements to hand on SKX serial square DGEMM with BLIS 0.7, OpenBLAS 0.3.10, and MKL mkl-2021.1-beta06 using the framework for the figures on the BLIS site:
size BLIS OpenBLAS MKL
2400 90.8 98.3 94.5
2352 91.4 97.4 93.6
2304 90.6 98.3 93.9What's even worse in real-world applications is that OpenBLAS misbehaves when an application uses threads. This is also described in the OpenBLAS FAQ:
If your application is already multi-threaded, it will conflict with OpenBLAS multi-threading. Thus, you must set OpenBLAS to use single thread as following.
The reason it's serial BLAS that mainly matters is that HPC codes are usually parallelized above the BLAS; why do you want the nesting? Swapping in threaded OpenBLAS or BLIS is something you might do with basically serial stuff like vanilla R, e.g. https://loveshack.fedorapeople.org/blas-subversion.html#_add... OpenBLAS threading has been somewhat buggy, but the main problem with its OpenMP support currently seems to be that using OMP_PLACES kills it.
If you’re in a commercial setting, your company might have 1 cluster and N simulation programs.
If you are publishing an application, I still recommend using the intel_dispatch_patch.zip.