Eigen: A C++ template library for linear algebra
eigen.tuxfamily.org
eigen.tuxfamily.org
It’s less than ideal for small things when the size is known at compile-time. If one knows SIMD intrinsics, in some of these cases the Eigen’s implementation can be outperformed by a large factor like 2-4. Also it’s very hard to mess with RAM layout of some things (like sparse matrices), just too many layers of abstraction and too much template metaprogramming.
But still, out of the box the usability of Eigen is awesome. And until the code is written, debugged, integrated and benchmarked, it’s generally impossible to tell whether a particular algorithm gonna be a performance bottleneck. That’s why I’m mostly happy with the library.
https://stackoverflow.com/questions/58071344/is-eigen-slow-a...
For use cases where FP32 precision is enough, I usually use DirectXMath library https://github.com/Microsoft/DirectXMath for that. That thing is cross-platform in practice. Even when building things for ARM Linux, it’s easy to copy-paste required pieces, NEON support is there.
When I need FP64 precision on PCs, I usually proceed without libraries, using AVX intrinsics.
One advantage of Eigen's approach that I haven't seen mentioned here is that its templated design makes it easy to substitute custom scalar types for operations, which helps enable straightforward automatic differentiation and other such tools (e.g. I'm currently using Eigen to make a tracing JIT for computations like FK, etc. over scenegraphs).
[0] https://bitbucket.org/blaze-lib/blaze/src/master/
[1] https://bitbucket.org/blaze-lib/blaze/issues/266/blaze-vs-ei...
Their current performance page is here [0].
[0]: https://eigen.tuxfamily.org/index.php?title=Performance_moni...
What's "RAM placement" in this context?
On the other hand, Eigen didn't claim to be faster than any BLAS during my interaction with informational side of their website. Nowadays, I just get the latest version, update my code and directly reach for their reference guide.
However, being used by both CERN (in LHC ATLAS) and TensorFlow, I guess they're no slouch in terms of performance (it also genuinely amazed my heavy BLAS fan professor at my Jury when the performance numbers came up). Another plus is, Eigen's ability to be used in CUDA kernels (though I didn't try it yet).
On the other hand, they're insanely fast (almost BLAS fast in my experience) for what I'm doing, and it's really well optimized (shows when you use -O3, especially on the solvers side). Also, it has neat capabilities like dynamic matrices and instant submatrix access (you can get a submatrix of a matrix, or just patch in a submatrix in a single call).
About RAM placement: From my experience, code becomes overly dependent on non-fragmented memory when allocating big matrices and/or vectors naively (i.e. in a contiguous manner), and execution fails as the node you're running on (or your computer) gets busier with other code. Eigen counters this by intelligently dividing bigger matrices into smaller chunks, both optimizing the access time and reducing (or effectively eliminating from my experience) the risk of segmentation faults related to failed memory allocations due to not having not enough undivided space in the memory space.
It claims to be faster than any free BLAS in the FAQ (but talks about ATLAS and GotoBLAS there). Given how close GotoBLAS in the guise of OpenBLAS gets to peak performance for GEMM, that would be dubious without measuring 2/3 of OB DGEMM performance. I assume claims of speed are for when BLAS is actually providing it. (I'm too experienced to be susceptible to appeals to authority of HEP software!)
Thanks for explaining the memory thing. I'd hope HPC code wouldn't be subject to SEGVs due to memory allocation failing (for two reasons), but you typically want arrays on the stack and otherwise large pages to minimize TLB misses (per Goto).
I'm not intending to bash Eigen, happy if people can use performant free software.
The only major issue, DirectXMath doesn’t support FP64 precision.
About the template stuff, I think their main performance advantage over traditional BLAS is not even SIMD, it’s lazy evaluation. Expressions like x=a*b+c never compute the complete a*b matrix or vector. The a*b expression returns a small placeholder object on the stack, of a scary type with couple lines of template arguments in the type name. This way the complete expression runs without making temporary matrices, instead it streams data from all 3 arguments and only writes to memory once. And if the `x` is of the correct size already, it doesn’t call malloc/free.
I'm sure someone has already looked at this, but I wonder if Eigen can just somehow nab kernels directly from BLIS, haha.
Eigen makes extensive use of expression templates in C++ to collapse complex operation sequences into streamlined and minimal calculations. This is generally ok, until you need a debug build. I've regularly seen debug builds of software using Eigen run 1000x to 10000x slower than the release build, which seriously complicates various debugging workflows. It also makes it a nightmare to run your test suite through valgrind, for example.
I've seen several engineers attempt (and fail) to try creating/linking a release build of Eigen with a debug build of the rest of the program to try to regain most of that speed while still allowing a decent amount of debugability, but this is really hard due to all the aggressive inlining and heavy use of templates.
In my experience, I would happily accept a 2x or more slowdown in linear algebra performance in release builds in exchange for significant boost in debug execution speed. If you're starting a greenfield project, you should consider how important decent debug performance is before choosing Eigen by default.
In debug mode it has no notion of performance whatsoever.
Running anything through valgrind or cachegrind will have several orders of magnitude slowdown - that's inherent to how the tools work.
To your last point, just add -g to your compile flags and see how far you get.
For example, we had simulation test suites which would run ~30 seconds of real-world time. If we can run those sims at 10x real time, each only takes 3 seconds, great! But with optimizations turned off, heavy Eigen code would run 1000x slower, which means the sims now take nearly 3000 seconds to run when bug hunting in the non-Eigen related surrounding logic. This already makes the test suite virtually un-runnable in debug builds, but if you want to then run those binaries through Valgrind (which tacks on another 100-1000x performance tax) you're now talking about something which is completely intractable. You can't run optimized builds through Valgrind without potentially introducing false positives, so you can't even run Valgrind on only the optimized builds... you just can't run it at all because your linear algebra library is slowing everything down.
The core problem is Eigen's heavy use of expression templates, which the optimizer cuts through just fine. But without the optimizer turned on, it adds layers and layers and layers of unnecessary abstraction. Because these are all templates, C++ expands all this code at every call site, making it substantially less efficient and impossible to shim out by linking with an optimized build of Eigen itself.
The only approach I've seen get close to making this work was an attempt by an engineer to use explicit template instantiation to create a separate binary with all the necessary Eigen functions which could be compiled separately with the optimizer turned on. However, this proved too nightmarish because since Eigen can template on the dimensions of everything, you need to know up front exactly all the various instantiations of Eigen types the program uses. This might be doable by writing a clang plugin or something, but doing it by hand was just way too much work because the codebase was huge, so we couldn't justify spending more time on it.
https://eigen.tuxfamily.org/index.php?title=Main_Page#Licens...
> Note that currently, a few features rely on third-party code licensed under the LGPL: constrained_cg. Such features can be explicitly disabled by compiling with the EIGEN_MPL2_ONLY preprocessor symbol defined. Furthermore, Eigen provides interface classes for various third-party libraries (usually recognizable by the <Eigen/*Support> header name). Of course you have to mind the license of the so-included library when using them.
> Virtually any software may use Eigen. For example, closed-source software may use Eigen without having to disclose its own source code. Many proprietary and closed-source software projects are using Eigen right now, as well as many BSD-licensed projects.
Otherwise, we've seen a 8-10x performance increase in SolveSpace (CAD) in some situations after switching from home-grown matrix operations to Eigen.
I tested all this very heavily before Eigen had AVX-512 support. In that environment there might be some differences and I would suggest you benchmark both configurations.
Hyperthreading does in general share units (both ALU and others); that's what hyperthreading is.
Apart from that, it really depends on what operations you're doing; e.g., modern Intel CPUs have three ports that can issue a 256-bit FMA, each, every cycle.
Yes, I think the issue with Eigen is cache related. They apparently have optimizations that are aware of cache architecture and running 2 threads that share the same cache will screw that up, resulting in more misses. If this is the case, I'd prefer algorithms that are cache line size agnostic. It is still much faster than the simple hand-written code we had before!
const float residual = (L.array() * P.array()).colwise().sum().square().mean();
const float residual = L.cwiseProduct(P).array().colwise().sum().array().square().mean();
const float residual = (L.transpose() * P).diagonal().array().square().mean();
The compiler can optimize using static information, e.g. these would all be handled differently for the following types Eigen::Matrix<float,3,3> L(3,3), P(3,3);
Eigen::Matrix<float,3,Eigen::Dynamic> L(3,K), P(3,K);
Eigen::Matrix<float,Eigen::Dynamic,Eigen::Dynamic> L(N,K), P(N,K);* If a vector is an n-by-1 array (a "matrix"), why not an n-by-1-by-1-by-1-by-1?
* What is a row vector in something like a function space? Is sin(t) a row or column vector?
* You _can_ think of a matrix as acting "to the right (on a column vector)" and "to the left (on a row vector)", but this disrupts the notion of a matrix as a linear function acting on a thing (since it can act on two things). It's much more consistent to define a matrix acting on a vector, m(x), independent of "direction" and then define the transformations that let you go in the opposite direction (take the transpose of m which is a whole new matrix).
* Are physical quantities like velocity a row or column vector? Convention dictates they're column vectors, so in what scenario would it become a row vector? If I have a velocity, v1, that's a row vector and another, v2, that's a column vector, is it appropriate to calculate the projection of v1 onto v2? Should the projection operation be "aware" of the orientations? If all vectors are the same ("column vectors"), this isn't a problem -- you can always define the projection operation without special cases.
* In matlab specifically, if you have an n-by-1 array `a` and a 1-by-n array `b` with the same elements, then `a(i) == b(i)` but `a` and `b` are not the same thing.
Or equivalently https://jiahao.github.io/talk/2017-06-juliacon/
There are basically two self-consistent choices for the design of linear algebra libraries.
But tge nice part is that these vectors can still interact well with parts of the code that expect a matrix.
It makes me wish for a language + compiler where linear algebraic objects are first-class values.
Like Fortran?
A similar one to consider that can at times be slightly easier to use coming from a python background is armadillo: http://arma.sourceforge.net/
The max performance was in Eigen-calling Intel MKL, but it was a big plus to not need MKL licenses on every development machine.
What did I do wrong?
It wraps Vulkan in a very thin layer, mostly to eliminate boilerplate.
Go
Category error.
One is a standard the other a library. BLAS covers a subset of what Eigen does, and Eigen can use BLAS/LAPACK routines directly from other libraries (e.g. MKL) for those things.