Blaze: High Performance Vector/Matrix Arithmetic Library For C++
bitbucket.org
bitbucket.org
Also: why is it that for projects like these, the very first 10 lines on the web site is not an actual code example of what usage looks like?
I had to actually go dig in the test directory to find actual code.
https://bitbucket.org/blaze-lib/blaze/src/master/blazetest/s...
Other than this, there really isn't much more to pick from.
Fair enough, but still doesn't tell me what's different about this project.
Saying BLAS is a C library skips over much of the depth and history of its FORTRAN roots. The work that went into achieving numeric stability of this stuff is astounding. Ditto with LAPACK.
https://en.m.wikipedia.org/wiki/Basic_Linear_Algebra_Subprog...
There are maybe one hundred libraries in the field.
IMO all of them and certainly any new/active one(s) should clearly differentiate (or state) thier purpose/goals and what makes them different. Please note, I am not saying what makes them better but different.
I am also tired of reading/hearing "HPC/simd/parallel" without ANY benchmarks/timings to support such claims. Isn't it strange that software implementing mathematics make claims (almost all of the times) without any measurements/proof?
Thanks for the info though, I might remember to try it in the future.
edit: the rant is not about Blaze, it is more of a general rant about similar libraries. I feel I have to mention that Blaze at least tries to address what I am complaining and I also saw that they have instructions on how to replicate their benchmarks on ones' machine which IMO is what every project should be doing.
>> Blaze is certainly more tuned for performance and uses nefty techniques under the hood like padding, explicit loop-unrolling and turning divisions to multiplications to create more opportunities for FMA. However, aggressively tuning for performance implies more specialisations, more overloads, more “SFINAE” and so on and as result compiles slower than Eigen.
To be clear, I think it’s great to have multiple competitive C++ implementations out there.
Eigen's internal structures and features doesn't map well to traditional BLAS/LAPACK structures anyway.
[0]: https://eigen.tuxfamily.org/dox/TopicUsingBlasLapack.html
BLAS/LAPACK is very optimized and it's recommended to compile on (or for) target for best performance, however they're developed for a very long time, and somewhat old-fashioned in terms of ergonomics and internal flow.
OTOH, Eigen is very modern, very easy to optimize (just pass -O3, and relevant -march -mtune to gcc), and you're screaming at 98% speed of BLAS/LAPACK.
I've used it extensively in my Ph.D., and TensorFlow is also using Eigen. It's very easy and practical to use, and it's very very fast. It makes abusing (ehrm making full use of) your processor easy and strangely enjoyable.
Some older benchmarks and current performance monitoring pages can be found at https://eigen.tuxfamily.org/index.php?title=Benchmark
HOWEVER, the lack of built-in GPU support is severely limiting, and we found the third-party blaze_cuda library very difficult to get to work. Since supercomputers are increasingly accelerator-focused, this limitation a big enough impedance that we made the extremely nontrivial decision to start fresh, on top of pytorch (where we know that there are years and years of development and support for accelerators ahead), instead.
Other than that, there are libraries that allow you to call CUDA from Java, like TVM and TornadoVM.