Intel Distribution for Python
software.intel.com
software.intel.com
I’ve been a little shy about using intel software since reading about this years ago
Edit, comparison:
$ perf record target/release/gemm-benchmark -d 1024
Threads: 1
Iterations per thread: 1000
Matrix shape: 1024 x 1024
GFLOPS/s: 96.36
$ perf report --stdio -q | head -n3
97.18% gemm-benchmark gemm-benchmark [.] mkl_blas_def_sgemm_kernel_0_zen
1.94% gemm-benchmark gemm-benchmark [.] mkl_blas_def_sgemm_scopy_down16_bdz
0.78% gemm-benchmark gemm-benchmark [.] mkl_blas_def_sgemm_scopy_right4_bdz
After disabling Intel CPU detection: $ perf record target/release/gemm-benchmark -d 1024
Threads: 1
Iterations per thread: 1000
Matrix shape: 1024 x 1024
GFLOPS/s: 129.12
$ perf report --stdio -q | head -n3
97.02% gemm-benchmark libmkl_avx2.so.1 [.] mkl_blas_avx2_sgemm_kernel_0
1.77% gemm-benchmark libmkl_avx2.so.1 [.] mkl_blas_avx2_sgemm_scopy_down24_ea
1.02% gemm-benchmark libmkl_avx2.so.1 [.] mkl_blas_avx2_sgemm_scopy_right4_ea
Benchmarked using https://github.com/danieldk/gemm-benchmark and oneMKL 2021.3.0.I'm really surprised popular numerical computing Python packages don't already have optimized hardware back-ends for things like NumPy... similar to ORC (OIL) which has been around for quite some time:
https://github.com/GStreamer/orc
But I don't know that much about Python under the hood, and I'm willing to be since so many academics work on this there's already optimized FFIs. I've used TensorFlow and it can offload tensor math to GPUs, but only NVIDIA's AFAIK.
I think it is hard to beat modern BLAS implementations for common operations. E.g. Apple Accelerate (which also implement the BLAS/LAPACK APIs) uses undocumented AMX instructions for large speedups compared to an ARM NEON implementation.
If there are true optimizations to be had, wonderful. But those should be added to core binaries pypi / conda. I am worried that Intel here may be trying to again artificially segment their optimization work on their math libraries for business rather than technical reasons.
They have.
PyPI:
https://pypi.org/user/Intel-Python
https://pypi.org/user/IntelAutomationEngineering
Anaconda:
When this first cropped up, I was using Digg.
Edit: removed note that they fixed the cripple AMD function; they didn’t, they actually just removed the workaround that made it easier to disable the checks; I was misinformed. Apparently now some software does runtime patching to fix it, including Matlab...
echo "deb https://apt.repos.intel.com/oneapi all main" | sudo tee /etc/apt/sources.list.d/oneAPI.list
You can read the "apt" section of the package managers, if that's what you prefer. https://software.intel.com/content/www/us/en/develop/documen...I once was excited about Intel releasing their own Linux distro (Clear Linux), but it has the same problem. It looks like Intel is trying to make custom optimized versions of popular open-source projects just to get people to use their CPUs, as they lose their leadership in hardware.
I don't know if there's an ARM-specific equivalent, but, if you want to use TensorFlow or PyTorch or whatever on ARM, they'll work quite happily with the Free Software implementations of BLAS & friends. If you code at an appropriately high level, the nice thing about these libraries is that you get to have vendor-specific optimizations without having to code against vendor-specific APIs. Which is great. I sincerely wish I had that for the vector-optimized code I was writing 20 years ago. In any case, if ARM Holdings or a licensee wants to code up their own optimized libraries that speak the same standard APIs (and assuming they haven't already), that would be awesome, too. The more the merrier. How about we all get in on the vendor-optimized libraries for standard APIs bandwagon. Who doesn't want all the vendor-specific optimizations without all the vendor lock-in?
Alternatively, if you would rather get really good and locked in to a specific vendor, you could opt instead to spam the CUDA button. That's a popular (and, as far as I'm concerned, valid, if not necessarily suited to my personal taste) option, too.
Plus, who's surprised? This is how Intel makes money. The consumer segment is a plaything for them, the real high-rollers are in the server segment, where they butter them up with fancy technology and the finest digital linens. Is it dumb? A little, but it's hardly a "problem" unless you intended to ship this software on first-party hardware which, hint-hint, the license forbids in the first place.
At the end of the day, this doesn't really irk me. I can buy a compatible processor for less than $50, that's accessible enough.
I think this is more likely aimed at AMD than Arm - don't think Arm is yet a threat in this space - and whilst they're entitled to do what they want it does make me less enthused about Intel and frankly more likely to support their competitors.
I'm not sure it's a sin for hardware manufacturers to support their products? In the days of yore, we even expected it of them.
I may be wrong but my experience is that AMD has been a bit better on this is the past e.g their OpenCL libraries supported both Intel and AMD whereas Intel's were Intel only.
AMD, on the other hand, also supplies Radeon GPUs for use with Intel CPUs. For example, that's the setup in the computer on which I'm typing this.
So I have a hard time seeing anything nefarious there. The one is obviously a business necessity, while the other would obviously be silly. Perhaps that changes with the new Xe GPUs?
I can see how if you've invested a lot in software you'd like to get a competitive advantage over your nearest rival so maybe a price we have to pay.
EDIT: Ah, I found it, AVX2.
conda create -n intel -c intel intel::intelpython3_core
Or [2]: docker pull intelpython/intelpython3_core
Note that it is quite bloated but includes many high-quality libraries.You can think of it as a recompilation in addition to a collection of patches to make use of their proprietary libraries.
Other useful links to reduce the noise in this thread: [3], [4], [5], [6].
[1] https://software.intel.com/content/www/us/en/develop/article...
[2] https://software.intel.com/content/www/us/en/develop/article...
[3] https://www.nersc.gov/assets/Uploads/IntelPython-NERSC.pdf
[4] https://hub.docker.com/u/intelpython
For example: .... benchmarks with this python is XXX % higher than ... (std python, AMD, ARM)But, 6 years ago, when I was in grad school, just swapping to the Intel build of numpy was an instant ~10x speedup in the machine learning pipeline I was working on at the time.
No idea if that's typical or specific to what I was doing at the time. I don't use MKL anymore because ops doesn't want to deal with it and the standard packages are already plenty good enough for what I'm doing nowadays. If you forced me to guess, I guess I'd have to guess that my experience was atypical.
https://en.wikipedia.org/wiki/Advanced_Vector_Extensions#CPU...
You could write your distro to check for flags that will tell it whether or not you have these using flags from /proc/cpuinfo. Or you could check whether it's in the Intel half of the list or the AMD half of the list. Or you could write your own distro that only runs on the first half of the list.
I get that Intel's contributions aren't purely altruistic. There are likely to be subtle tuning problems that require slight changes to optimize on different platforms, and they can't really be expected to do free work for AMD. But it looks to me like they're being unecessarily anticompetitive.
Isn't setting up barriers to entry generally considered to be a part of healthy competition? I'd hazard to say that as long as a company is playing within the boundaries of what's allowed, there's nothing they could do that's anticompetitive; at the most, you could accuse them of being somewhat unsportsmanlike.
No, it is not. This is better described as vendor lock-in, than a barrier to entry. But vendor lock-in is also against healthy competition.
Healthy competition means that users choose your product because it suits their needs the best, not because they are somehow forced to choose your product.
Artificial barriers to entry are contrary to that and if they’re not illegal they should be.
I fully admit that there are natural barriers that occur at times. I don't think that you should be expected to reverse-engineer your competitor's products and bend over backwards to make them work better.
Here, for a concrete example, Intel had a clear choice to test whether a processor supported a feature by checking a feature flag - It's in the name, they're literally implemented for that exact purpose - or they could expend extra effort in building their own feature flag database by checking manufacturer and part number. They could have either expended extra effort to launch and distribute their own entire custom Python distribution, or submitted pull requests to the existing main distribution. For another example, Apple could have used industry-standard Phillips or Torx screws in their hardware: Manufacturers had lines to produce them, distributors had inventory of the fasteners, users had tools to turn them. Instead, they went to great expense to build their own incompatible tri-lobe screws, requiring probably millions of dollars in investment in custom tooling and production lines, all for the sake of creating an artificial barrier.
A 10% speed improvement on 1000's of jobs could in theory save you a nice chunk of time. This becomes very important in the financial market where you need batch jobs to be finished before markets open, or you just want to save 10% on your EC2 bill.
So they definitly don't agree you can do the same with free software.
However, I recently ran the Polyhedron Fortran benchmarks with the compilers to hand (~2020 vintage). XL (on POWER) was the only one that gave a significantly better bottom line; obviously IBM know how to compile Fortran well by now (or by Fortran-H). As far as I remember, that was essentially due to the treatment (vectorization) of maths intrinsics, which probably aren't so dominant in typical HPC code. One bad case -- fmod inlining -- has since been fixed in GCC. Without GCC's unfortunate longstanding failure to vectorize sincos (or equivalent), gfortran should have beaten ifort significantly in the bottom line, and at least got close to XLF. PGI was distinctly worse than GCC, but may do better at OpenACC/GPU offload, for instance. XL may win on OpenMP, since some of the current standard was for Sierra's needs. I should find time for the NAS benchmarks.
Thankfully, accelerated math libraries already exist for Python without the vendor lockin.
Why not use the original distributions?
One use might be improving throughput of a compute bound system, like an etl written in python, with little effort. Ideally just downloading the new interpreter.
I guess that is my naive wish list for a short term speed up :)
Thanks for the heads up!
EDIT: Oh god it's 1 indexed
It's obviously not relevant to Python per se, but you get basically equivalent performance to MKL with OpenBLAS or, perhaps, BLIS, possibly with libxsmm on x86. BLIS may do better on operations other than {s,d}gemm, and/or threaded, than OpenBLAS, but they're both generally competitive.