BLIS: A BLAS-like framework for basic linear algebra routines
github.com
github.com
It seems that the selling point is that BLIS does multi-core quite well. I am especially impressed that it does as well as the highly optimized Intel MKL on Intel CPUs.
I do not see the selling point of BLIS-specific APIs, though. The whole point of having an open BLAS API standard is that numerical libraries should be drop-in replaceable, so when a new library (such as BLIS here) comes along, one could just re-link the library and reap the performance gain immediately.
What is interesting is that numerical algebra work, by nature, is mostly embarrassingly parallel, so it should not be too difficult to write multi-core implementations. And yet, BLIS here performs so much better than some other industry-leading implementations on multi-core configurations. So the question is not why BLIS does so well; the question is why some other implementations do so poorly.
MKL is a general-purpose library, it does compromises to be good for most sizes. If you care about the performance of specific sizes, you can write code to beat it on that very specific benchmark.
Interestingly, BLIS does poorly at small sizes, which is the area where it is easiest to beat MKL at.
It is a specialized library and they JIT their kernel.
Usually BLIS is slower.
If I recall correctly, MKL also has interfaces that allow different array orders (row order or column order).
It's the exact opposite, most numerical linear algebra is _not_ embarrassingly parallel and requires quite an effort to code properly.
That is why BLAS/LAPACK is popular and there are few competing implementations.
My professional and academic background is in numerical analysis and scientific computing (including BLAS/LAPACK level implementations), but I admit I haven't done deep numerical implementations in years.
At the other end of the spectrum, getting a matrix-matrix multiply to run fast isn’t easy either. It’s what necessitated the kind of approach the authors of BLIS adopted. On paper it’s easy, but actually getting it to run fast on a computer isn’t.
With that out of the way, basic linear algebra operations that require sophisticated algorithms and are not "embarrassingly parallel": matrix multiplication, matrix inversion, matrix decomposition (SVM, QR, etc). Some of these algos fall into BLAS and others in LAPACK.
And then, there are many parts of the algorithms deep in things like the various matrix decompositions that are difficult to parallelise because of non-trivial data dependencies. It’s easy to write some code to pivot a matrix; it’s very hard to do it in an efficient and scalable way. And because all the higher-level routines depend on them, so they have a large effect on overall performance.
But from this exercise alone, knowing the tricks you used already, you can see how un-embarrassingly parallel this task is (frankly if it is truly embarrassingly parallel, then `#prama omp for` should bring you close to best possible performance already.)
I don’t think your prof got it wrong, it is just that you misunderstood them.
As soon as you start mixing numerical stability (data-dependant operation reordering) concerns into linear algebra routines -- arguably one of the main points of LAPACK -- parallelization necessarily gets trickier because you either need a good enough heuristic to ignore those reorderings or you also need data-dependant scheduling without too many pathologies.
https://danieldk.eu/Posts/2020-08-31-MKL-Zen
Some back history: they used to use (slow) SSE kernels when a non-Intel CPU is detected. Over time, they started adding kernels for Zen, but last time I tested the Zen kernels were still slower than the AVX kernels. So it still pays off override the Intel CPU detection.
It looks like these benchmarks use MKL 2020 update 3. If I recall correctly, this version did not have the Zen sgemm kernels yet. A newer MKL version would perform much better and if you disable Intel detection, MKL would be competitive on Zen.
Often you get tensors that are sliced views on one or more dimensions, with BLAS-api you need to allocate them in a contiguous buffer. BLIS does not need that, there is a repacking algorithm to overcome memory bandwidth bounds in matrix multiplication (search the "roofline model"), the slicing/realloc can be fused with repacking with BLIS API.
BLAS/CBLAS assumes your matrices are column-major/row-major. BLIS uses general stride instead (both columns and rows are parameterized by stride) which makes it more flexible on matrix storage and makes it easier to implement tensor contractions in higher dimensions.
This necessitates a different API.
https://www.edx.org/learn/computer-programming/the-universit...
Neat stuff. The instructors are great, so I'll try to make room in my schedule for this later in the year.
BLIS: BLAS-Like Library Instantiation Software Framework - https://news.ycombinator.com/item?id=11369541 - March 2016 (2 comments)
is there any documentation on the architecture and engineering of OpenBLAS kernels? As a non-expert, delving into the assembly code did not give me any insights on how it worked, and I would love it if there was some way to ease people like me into the deep operation of those routines. Particularly, I would love to understand the engineering against CPU specifics (caches, superscalar arch, speculative ex.) and similar low-level insights.
Thanks in advance!
It even has a nice, friendly licence.
ATLAS tries to automate it away.
For BLIS, instead they isolated the part that requires hand-tuning so you only have to write a couple kernels, then they can use them for the whole library.
BLIS is quite a bit newer and I’d be surprised if ATLAS kept up.
The interface isn't perfect though. For example they didn't commit to supporting zero strides. That used to work anyway, but it's broken now for some archs.
AMD did use it to create their BLAS library though.
Also, side note, when I’ve heard “reference BLAS,” it has been used in the opposite way. Netlib BLAS is the reference BLAS, it has basically bad performance but it defines the functionality. AMD used BLIS to create a tuned vendor BLAS.