Faster: Fast numerical calculations in Rust
github.com
github.com
Please consider adding such benchmarks so that we can assess how fast this library really is when compared to other more established libraries.
It shows four lines of code, and compares it to ~30 lines with the comment "Even with all of that boilerplate, this still only supports x86-64 machines with SSE or AVX ... "
The closing section: "Here are some extremely unscientific benchmarks which, at least, prove that this isn’t any worse than scalar iterators." suggests the final executed code probably runs at around the same speed, but the short version is _way_ easier to write and _way_ more portable.
having run into this while working on Rust code a while ago, I really appreciate this library! I would expect Rust users to come to this library when they find themselves working with numeric arrays on Intel CPUs, much like I did.
I think it would be quite useful to me right now, whether or not it turns out to be fast enough to make users doing array computations come to Rust. (I'd guess probably not, although it could get there!)
Right now, it's within 5% of ugly-style explicit SIMD code in the absolute worst case (the example in the README is more like 1%). The hot bit (the body of the map) compiles down to 100% simd instructions in the order LLVM thinks will be the fastest, and the load/store bit has an extra branch per load/store. This is because faster is mainly a bunch of 1-line wrappers around SIMD intrinsics, and a bunch of polyfills which only get used on machines which don't support those intrinsics.
The main reasons you'd experience a perf hit compared to explicit SIMD when using this library would be using runtime feature detection (minimal, and not implemented yet - would add around 2 branches per simd_iter call), getting lazy with gathers/scatters rather than ugly bit-level hacks, and accidentally using operations which aren't vectorizable on your machine (scatters on non-avx512, reductive operations on non-sse4.2, etc) and not realizing it (explicit simd will throw a SIGILL; faster will use a scalar polyfill).
This crate explicitly focuses on usability, and given the limited author time, prioritizes that.
lay off the unnecessary aggressiveness.
Neither did the programmer of this lib.
>Extraordinary claims require extraordinary evidence.
Only if one's goal is to convince people who can't be bothered to do the check for themselves. Otherwise one can just make the claim, and let others (if they are interested) sort out whether it's true.
On my C code, I have a build script that rebuilds a shared library 3-4 times for different CPUs (e.g. with and without -mavx2) and then uses dlopen() to choose the right one at runtime. The actual SIMD code has a bunch of #ifdefs to choose the appropriate implementation for primitive operations.
GCC has function multiversioning [0][1] (__attribute__((target("sse4.2")))) to enable choosing a function implementation based on CPU capabilities at runtime, but that is not suitable for the use case I describe. I want my "high level" code to be the same, but using the right implementation for "primitive" ops like dot product. Using multiversioning would completely kill the optimizer and cause excessive register spilling and other unwanted artifacts.
[0] https://gcc.gnu.org/wiki/FunctionMultiVersioning [1] https://gcc.gnu.org/onlinedocs/gcc-6.2.0/gcc/Function-Multiv...
Yes, You are right! I understood, that they wanted to add the branch to _every_ SIMD instruction. I concur with branching outside of loops.
Re Rust: I expect that one could use a macro around the function (or annotate it, if the macro system is good enough) that generates these different functions and branches to the right one.
Core is https://github.com/ekmett/rts/blob/master/src/rts/varying.hp... and https://github.com/ekmett/rts/blob/master/src/rts/vec.hpp .
[1] eg, random ggl link: https://gain-performance.com/2017/05/14/umesimd-tutorial-8-c...
https://github.com/AdamNiederer/faster https://github.com/rayon-rs/rayon
I look forward to learning more Rust so I can get an excuse to use this.
That's actually awesome! CPU's and GPU's are capable of breakneck speeds, so I love to see things that really make the most out of the capacity available.
STOP
C Add a couple more stop statements in case someone runs this on a Cray and it is going so fast it breaks though the first one:
STOP
STOPYes, Linux does the equivalent of
while(true)
{
halt
}
I don’t think that code will loop, but why take the risk?There is/was a detailed description of what it takes to stop an OS that included multiple ways to stop the CPUs, executed in order in the hope/expectation that one of them would work that in the end hit a similar loop, but I can’t find it anymore.
When I was young I never understood why someone would turn turbo mode off.
Are we going to see a RustBLAS soon? Something open-source that could go up against MKL for performance would be really amazing.
But the point would be no longer having to trade software freedom (or at least open source) for performance -- try benchmarking Numpy built on MKL vs OpenBLAS and you'll see what I mean.
If someone gets inspired by a new language to try and write a free or open-source MKL competitor, because that language has such-and-such features that make it more feasible than in Fortran or C, then that'd be fantastic. All I was suggesting is that it'd be cool if Rust was that language.
Superlatives are usually challenged and you will require proof. The proof itself can also be challenged. Suddenly you are spending a lot of time.
I would rather say: "SIMD math library" rather than "Fast library".
I am just generally against referring to software in absolute terms like: simple, fast, friendly, lightweight, etc. since it can be misleading.
I consider a good practice to separate facts and opinions to achieve clear communication. Passing opinions as facts is dogmatism.
Except for autocomplete...
I know Java does this.
[1] - https://software.intel.com/sites/landingpage/IntrinsicsGuide...
A few examples I'm targeting for the next point release are matrix determinants, cross products, byte-aligned encoding (like https://github.com/AdamNiederer/base100/blob/master/src/lib....), and reductive operations like strcmp. AFAIK none of that can be autovectorized at all.
It's just pretty hard to create one. Same applies to automatic optimizations in the general case: what do you exactly know about what Java does beyond that it is capable of using SIMD instructions? Have you perhaps evaluated how and where it does so, or how the end result compares to more explicit manual use of stream operations?
edit: and for the record, the problem is not as easy as you seem to think it is, given how you described your astonishment.
Kinda-sorta yes, but not really. Auto-vectorization is a very brittle optimization that only works at the best of times and the slightest change in your code can cause it to fall out of the hot path. You have to be extra careful with aliasing and use the "restrict" keyword a lot.
If you really do care about performance, you still have to write your hot loops with explicit SIMD code.