Blaze: A high-performance C++ math library
bitbucket.org
bitbucket.org
https://bitbucket.org/blaze-lib/blaze/wiki/Benchmarks
Turns out all benchmarks are done circa 2016/17 and use pretty old versions of libraries and compilers. That's a long time. Any excitement I had was definitely tempered, would be good to have a recent update.
I gather Kompute.cc, which runs on GPUs via Vulkan, is the modern choice. Running floating point array kernels on CPU cores seems pointless these days.
(That said, I run QubesOS, and literally everything I run does all its work on CPU cores--including all Vulkan code--because nothing can see the GPU.)
However, in some cases, where GPU brings substantial performance benefits in a long running (hours, days, etc.) application, you can setup pinned memory locations, where they map to GPU. Then you can feed the input part, where it feeds the GPU. If you have multiple DMA engines in the GPU, you can just stream the data from in to GPU to out, and read the results from your second pinned memory location.
If the speed bump is big enough, and you can eat the first startup cost, PCIe cost can become negligible.
However, it’s always horses for courses, and it may not fit your case at all.
In my case, Eigen is already extremely fast for what I do on CPU, and my task doesn’t run long. So adding GPU doubles my wall clock.
But, as always, fast enough is fast enough.
Or you can just use Eigen via CUDA:
https://eigen.tuxfamily.org/dox/TopicCUDA.html
> Running floating point array kernels on CPU cores seems pointless these days.
CPU level vectorization can do wonders in terms of speed with right libraries (like Eigen, MKL, etc.). I can solve the same problem in seconds with Eigen where its MATLAB implementation takes hours.
I gather there are efforts to provide a conversion layer from CUDA to other targets, but they seems to involve adapting the CUDA code to be more portable.
While I use Eigen extensively on CPU, I don't follow any mailing lists of them, so I don't know the internal state and mindset.
The link you’ve posted doesn’t mean you can do large dense matmul in CUDA right away, it’s just that you can use fixed sized vector/matrix operations inside kernels (which might be useful if you’re writing graphics code and need some vec3/mat3/quats.) You still have to write your own customized kernel for a large dynamic-sized gemm computation (with tiling and shared memory and all that jazz), and at that point it’ll be best to just use CUBLAS.
> By default, when Eigen's headers are included within a .cu file compiled by nvcc most Eigen's functions and methods are prefixed by the device host keywords making them callable from both host and device code.
Eigen casually overrides "*" to do any kind multiplication, hence I guess it'll also carry the GeMM functions alongside to the kernel during compilation, but this needs to be tested.
Considering the speed I got from running Eigen on CPU, I still need to find larger problems to make that effort worthwhile, however.
This is true only for 32-bit or lower precision floating-point numbers.
For 64-bit or higher precision floating-point numbers, which are required for the bulk of the scientific computation applications, CPUs are much better than the modern GPUs.
In the latest generation of both NVIDIA and AMD "consumer" GPUs, the throughput ratio of FP64 vs. FP32 has been reduced to 1:64. This ensures that these GPUs have both a performance per dollar and a performance per watt that are worse than those of CPUs.
The "datacenter" GPUs have excellent performance per watt, but their immense prices ensure that these GPUs either have a performance per dollar that is worse than that of CPUs (except for those who buy hundreds or thousands, to be used 24/7, so the savings in power consumption can balance the high acquisition cost) or even when the performance is so high that the ratio performance/price is similar to CPUs, the magnitude of the price is so great that it is far outside the range that can be afforded by a small business or by an individual.
More than a decade ago, NVIDIA claimed continuously that the GPUs will completely replace the traditional CPUs for all high-volume computations. Nevertheless, they themselves are those who have prevented this from happening, by segmenting the GPU market and raising the price of the 64-bit GPUs by more than an order of magnitude.
Eigen has commit-based performance monitoring for some time, arewefastyet style: https://eigen.tuxfamily.org/index.php?title=Performance_moni...
I've used Eigen to play with 3K by 3K dense matrices and solve them in some cases, and it's not even blink an eye, neither in time nor in space department.
Blaze: High Performance Vector/Matrix Arithmetic Library For C++ - https://news.ycombinator.com/item?id=28493373 - Sept 2021 (28 comments)
Benchmarks for Blaze, A high-performance C++ math library - https://news.ycombinator.com/item?id=10117971 - Aug 2015 (30 comments)
Ideally one would like a library that checks at least all of the following (in no particular order):
* open source with a good community around it
* written in popular, modern, easy to use programming language
* offers complete functionality: linear algebra (including ndarrays), special functions etc. Numerical Recipes is a reasonable checklist of what is needed
* has an intuitive, "math-like" API that expresses well mathematical expressions
* makes optimal use of all available hardware (multi-core, GPU, cluster) without the need for (too much) specialized programming or tuning
* can integrate easily with other programming languages / frameworks as part of bigger projects
Those days there are quite a few candidates, from numpy/python, to armadillo,eigen,blaze,xla/C++ to various rust and julia libraries. What we don't have is an easy way to select :-). On the plus side, if the project is still in the early phases / not too big, it is relatively easy to try out and migrate. But a detailed and updated comparison table (e.g. in Wikipedia) would indeed be quite handy.
They all use or can use the same core libraries to do the actual computationally intensive parts (BLAS, LAPACK) so they mostly share the code where optimisations would make the most impact. So, at the end of the day, it's mostly about the API you like because performance-wise it's likely to be a toss up.
it gets even more complicated if you expand into efficient random number generation, graph processing etc. then it becomes important that the library is extensible with some plugin mechanism etc. as, ideally, you don't want your project to be a potpourri of different libraries / API's / dependencies
all in all there are some amazing projects out there but sometimes more options means it is harder to decide without spending serious time. I think this is one advantage of the python scipy stack: at the expense of a (possibly small) performance penalty you have one choice only but it is fairly complete and consistent
> C++
I haven't used blaze, but the readme sounds similar. We generate kernels using expression templates and have a similar syntax to matlab or python.
https://github.com/NVIDIA/MatX
If anything is missing we're happy to take feature requests.
I can recommend the papers, they're quite accessible and there is something very satisfying about the math tricks they came up with to build the polynomials.
They do seem to have added more rounding modes though, which presumably makes the library more generally useful.[1]
[0] https://people.cs.rutgers.edu/~sn349/rlibm/
[1] https://blog.sigplan.org/2022/04/28/one-polynomial-approxima...
https://romanpoya.medium.com/a-look-at-the-performance-of-ex...
(Hint, hint.)