SIMD in Rust
huonw.github.io
huonw.github.io
I would have liked to see benchmarks compared to non-simd rust too. Thankfully this is in the source code. Here's what I get from cargo bench, fwiw:
Running target/release/mandelbrot-aefed80dbc3f2841
running 2 tests
test mandel_naive ... bench: 802,072 ns/iter (+/- 25,106)
test mandel_simd4 ... bench: 235,853 ns/iter (+/- 6,374)
test result: ok. 0 passed; 0 failed; 0 ignored; 2 measured
Running target/release/matrix-58de8ccd4bd58dcd
running 6 tests
test inverse_naive ... bench: 4,967 ns/iter (+/- 251)
test inverse_simd4 ... bench: 1,984 ns/iter (+/- 94)
test multiply_naive ... bench: 2,226 ns/iter (+/- 26)
test multiply_simd4 ... bench: 897 ns/iter (+/- 32)
test transpose_naive ... bench: 627 ns/iter (+/- 16)
test transpose_simd4 ... bench: 361 ns/iter (+/- 7)
test result: ok. 0 passed; 0 failed; 0 ignored; 6 measuredBTW, the benchmarks are comparing to non-SIMD rust: the graphs are of how many times faster the SIMD Rust code is than scalar Rust code (i.e. if they were plotted, the scalar bars would all be at 1.0).
This is in extreme alpha stages and is written by people who are still coming to grips with rust. If anyone wants to chat about it or has some feedback, please drop by over here: https://gitter.im/arrayfire/arrayfire-rust
But the results are already impressive; the team said that Servo provides a 2x speed increase when rendering the CNN homepage and a 3x speed-up on Reddit.
> Speaking of the optimiser, rustc uses LLVM, which is industrial
> strength, and supports a lot of autovectorisationIn any case, I agree the Mandelbrot example isn't so interesting: I included it because it is relatively simple, well-known and gives a pretty picture (i.e. good for a blog post where a single example isn't mean to be the focus). In fact, manual unrolling catering to autovectorisation is how Rust is currently top of the mandelbrot benchmark game[1], and explains the equal performance of the explicit-SIMD and scalar versions of spectral-norm on AArch64 (although the fact they aren't equal on x86 hints at the lack of guarantees around autovectorisation).
I find the examples like matrix inversion, nbody and fannkuch-redux are more compelling because the vectorised version is far less similar to the scalar one ("strange" shuffles, approximation of floating point ops and dynamic byte shuffles with precomputed values, respectively).
[1]: http://benchmarksgame.alioth.debian.org/u64/performance.php?...
How well does it work in general? When you write SIMD code, can the compiler keep the values in vector registers or is there spilling going on?
I'm planning follow up posts which may involve more assembly/IR, but this is designed to be an introduction/high-level post, and the graphs are meant to serve as a summary/replacement for digging through reems of assembly.
The bounds checking isn't that bad in and of itself, e.g. for f32x4 it is one bounds check for 4 elements and that bounds check is just a comparison and an extremely well-predicted branch. However, it definitely can be noticable and can inhibit other optimisations.
There's various routes this can be improved for sure, e.g.
- `unsafe` versions of the `load` that don't bounds-check and/or do aligned loads,
- functions that convert a `&[f32]` into a `&[f32x4]` (possibly with prefix and suffix &[f32]'s for unaligned left-overs),
- tweaking the set of optimisation passes the compiler runs to handle the patterns that occur in Rust better (I believe rustc just runs the default set of LLVM passes, which are likely more tuned to C/C++ than Rust, e.g. the IRCE pass[2] should help eliminate more bounds checks but isn't enabled by default because presumably C/C++ don't use bounds checks enough for it to be worth it)
- higher-level combinators/"algorithms" that avoid the need to do the loading/memory-management manually
[1]: http://huonw.github.io/simd/simd/struct.f32x4.html#method.lo...
However, there's no fundamental reason Rust can't be a viable replacement language for C for these sort of things. As one example, I think it would be great to have Rust used more in the numeric/scientific computing space, where a lot of code is written by non-expert programmers (i.e. don't mix well with C/C++ to get reliable software), so having a compiler looking out is possibly nice.
Also, it should be possible to replace a C component with a Rust one in essentially a drop-in zero-overhead way since Rust can trivially expose a C-compatible interface.
This isn't the way to go. Once you have to call an external function, you lose all compiler optimizations that might take place. If you write low level "primitive" operations (e.g. matrix multiply and inverse), you will get worse performance by doing external function calls rather than having your code made visible to the compiler so that it can be inlined and optimized.
Additionally, you can't write assembler SIMD code that is portable from architecture to architecture. This can be done in C using extensions but not if you stick to the standard (and support MSVC). Quite a lot of SIMD code can be written portably without having to use architecture specific instructions.
It's definitely a good thing that languages like Rust have native SIMD.