Faer-rs: Linear algebra foundation for Rust
github.com
github.com
Does anyone know how much of this is happening in the matrix/array space in rust? There are several libraries that have overlapping goals: ndarray, nalgebra, etc. How much do they share in terms of underlying code? Do they share data structures, or is anything like that on the horizon?
Good to see LU decomposition with full pivoting being implemented here (which is missing from BLAS/LAPACK). This gives a fast, numerically stable way to compute the rank of a matrix (with a basis of the kernel and image spaces). Details: https://www.heinrichhartmann.com/posts/2021-03-08-rank-decom....
n faer mkl openblas
1024 27.06 ms 186.33 ms 793.26 ms
1536 73.57 ms 605.71 ms 2.65 s
2048 280.74 ms 1.53 s 8.99 s
2560 867.15 ms 3.31 s 17.06 s
3072 1.87 s 6.13 s 55.21 s
3584 3.42 s 10.18 s 71.56 s
4096 6.11 s 15.70 s 168.88 s
I will agree with you on the routines being broken. That happened within the past year. Not sure why it happened, but it's annoying. If you know the name of the routine you can usually find documentation on another site.
>The BLAS-like Library Instantiation Software (BLIS) framework is a new infrastructure for rapidly instantiating Basic Linear Algebra Subprograms (BLAS) functionality. Its fundamental innovation is that virtually all computation within level-2 (matrix-vector) and level-3 (matrix-matrix) BLAS operations can be expressed and optimized in terms of very simple kernels.
Eigen handle most (if not all, I just skimmed the tables) tasks in parallel [0]. Plus, it has hand-tuned SIMD code inside, so it needs "-march=native -mtune=native -O3" to make it "full send".
Some solvers' speed change more than 3x with "-O3", to begin with.
This is the Eigen benchmark file [1].
[0]: https://eigen.tuxfamily.org/dox/TopicMultiThreading.html
[1]: https://github.com/sarah-ek/faer-rs/blob/main/faer-bench/eig...
Did you check with resource utilization? If you don't provide "OMP_NUM_THREADS=n", Eigen doesn't auto-parallelize by default.
My complete Ph.D., using ton of Eigen components plus other libraries was compiling in 10 seconds flat on a way older computer. This requires gigs of RAM plus a minute.
Top-tier numerical linear algebra libraries hold all hit the same number (give or take a few percent) for matrix multiply, because they're all achieving the same hardware peak performance.
i plan to address that after refactoring the benches by executing each library individually
Also stops them from grumbling after you post good results!
and of course i do check the cpu utilization to make sure that all threads are spinning for multithreaded benchmarks, and occasionally check the assembly of the hot loops to make sure that the libraries were built properly and are dispatching to the right code. (avx2, avx512, etc) so overall i try to take it seriously and i'll give credit where credit is due when it turns out another library is faster
overall, the results show that faer is usually faster, or even with openblas, and slower than mkl on my desktop
[0] https://github.com/sarah-ek/faer-rs/blob/main/faer-bench/src...
[1]: https://github.com/sarah-ek/faer-rs/blob/172f651fafe625a2534...
fn sum_of_squares(input: &[i32]) -> i32 { input.iter() .map(|&i| i * i) .sum() }
If you changed it to input.par_iter() then it executes in parallel on multiple threads.
nalgebra doesn't use blocking, so decompositions are handled one column (or row) at a time. this is great for small matrices, but scales poorly for larger ones
eigen uses blocking for most decompositions, other than the eigendecomposition, but they don't have a proper threading framework. the only operation that is properly multithreaded is matrix multiplication using openmp (and the unstable tensor module using a custom thread pool)