It's Time to Pay Attention to Intel's Clear Linux OS Project
forbes.com
forbes.com
The optimizations they are applying only exist in hardware made in the last ~9 years.
Nothing stopping any other distro from copying the approach other than detail work and increased package sizes.
So, why aren't these optimizations available on other distros, since they seem to make such a difference? Is it only a matter of time?
Yes, with these tests showing some ~10% improvements, some distros getting regularly spanked in benchmarks probably will.
It seems with this distro you won’t have to wait. But I can’t imagine the speed ups will be that noticeable.
Otherwise, Forbes article for this kind of "opinion" makes me very conscious about if this was "prompted" or not. No harm intended just want to be conscious about the behind the scenes stuff.
Author is Jason Evangelho (https://twitter.com/killyourfm) who seems pretty neutral.
For instance, https://clearlinux.org/news-blogs/transparent-use-library-pa... talks about architecture-specific versions of OpenBLAS, but on x86_64 OpenBLAS dispatches on the architecture. Also it has similar DGEMM performance to MKL, at least on for AVX2 and below, unless avx512 has been improved recently; BLIS is competitive also for AVX512. The example does imply something useful that they seem to have done. That's to add SIMD hwcaps for dynamic loading. This is an important omission from vanilla Linux/ld.so, which means you can't build specific libraries and automatically get the appropriate one loaded for the architecture you're on. (Obviously you can arrange to get the appropriate one with LD_LIBRARY_PATH or ld.so.conf, but...) As far as I remember, the only thing that works for on x86_64 is TLS, i.e. there's a /usr/lib64/tls on Fedora-ish systems, but not a similar one for avx2.
Clear Linux has a script which looks at GCC optimization reports and adds target_clones attributes to C(++?) functions which report they're vectorizable. Many of those won't be helpful. The example I've seen, but don't have to hand, was for FFTW but, like OpenBLAS, that dispatches to SIMD-specific kernels. What the script picks up in the example is useless; I think it's just in a test harness, but at least not something that will make your FFTs go faster. That sort of thing can actually be harmful if firing up the SIMD unit lowers the clock rate to no good effect.
There may be problems with that anyway. I haven't had a chance to investigate closely, but adding target_clones to the generic C kernels for BLIS' DGEMM doesn't get the performance that it should, compared with a straight -march=. [It might be worth noting that you can get about 2/3 the performance of the hand-tuned DGEMM kernels with the generic C and appropriate GCC flags.]
This stuff isn't specific to Intel hardware (v. AMD) except insofar as they choose specific targets, and Zen is similar to Haswell for linear algebra kernels.
Probably the biggest blame is Nvidia for making a compatibility mess with their drivers.
This would be for cuda on servers.