The state of SIMD in Rust in 2026
shnatsel.github.io
shnatsel.github.io
The ARM-based CPU manufacturers make this worse by posting almost no low-level documentation for their CPUs. For basically any mainstream x86 CPU, it's trivial to find documentation listing what ISA level it supports and general execution widths and latencies for common operations. For the majority of ARM CPUs, there's absolutely nothing. ARM only has optimization guides for selected Cortex cores, and NVIDIA published info for their Olympus core. But execution details had to be reverse engineered for Apple M1, and there is nothing for Oryon. This is especially bad for in-order cores, which unfortunately is still relevant because new CPUs are still being shipped with in-order efficiency cores.
Or LSE, adding a whole different set of atomic ops... which are way faster on some CPUs...
Are those not based on standard ARM cores or ARM documentation on its cores not deep enough for your purpose?
https://support.arm.com/documentation/uan0016/a/
This has detailed information on latencies and throughput, which are important when optimizing SIMD code. But ARM doesn't publish optimization guides for all their cores.
On top of that, the cores are often modified in significant ways. Snapdragon CPUs, for instance, have used modified Cortex cores in the past and can have performance differences from the original core.
To be fair, Intel's been slacking a lot on this too, not even bothering to update their own optimization guide for their latest cores. But that's made up for the community mining this information in great detail on sites like uops.info, and also being a lot less x86 cores to deal with. On the ARM side, there's practically not much analogous other than the Apple M1 microarchitecture analysis.
> AVX-512, on the other hand, has been moving backwards since Intel has not shipped any consumer-level CPUs with it for years now.
Good thing Intel hasn't shipped a viable consumer-level CPU in years either. What should never have happened is Intel removing AVX-512.
> Steam Hardware Survey
While it's a good source, it's not a definitive source. Four out of five amd64 machines at my place support AVX-512F, but none of them participate in that survey. And the one machine that doesn't have it... I wish it did.
At one point, Intel integrated graphics were literally the most common individual GPU in the Steam Hardware Survey - Intel HD 4000 held the #1 spot until the GTX 970 overtook it in late 2015. Does that mean developers shouldn't have targeted discrete GPUs either, just because most survey participants didn't have the hardware they wanted to target?
If the issue were just older hardware, we could just wait out and let people upgrade a bit. But new hardware are shipping without AVX512 unfortunately
No one's saying it shouldn't exist, but it's a different story to ship software that requires it unconditionally, particularly if you are targeting the consumer market instead of only servers.
Not sure why you keep referring to Skylake since Intel hasn't shipped a Skylake-based CPU in years, and the most problematic CPUs currently shipping that lack AVX-512 support are several architectural generations ahead of Skylake, for both the P-cores and E-cores.
> Good thing Intel hasn't shipped a viable consumer-level CPU in years either. What should never have happened is Intel removing AVX-512.
Agreed on not removing, but saying that Intel hasn't shipped a viable consumer-level CPU in years is silly. They still ship huge volumes of CPUs, especially in laptops where AMD is still underrepresented.
> At one point, Intel integrated graphics were literally the most common individual GPU in the Steam Hardware Survey - Intel HD 4000 held the #1 spot until the GTX 970 overtook it in late 2015. Does that mean developers shouldn't have targeted discrete GPUs either, just because most survey participants didn't have the hardware they wanted to target?
Targeting a discrete GPU is different than not supporting integrated GPUs at all, which would be analogous. And many games do have to support integrated graphics even if the performance is not great, precisely because they don't aim high enough in the market to be able to ignore iGPUs entirely.
I also mention the Steam Hardware Survey because it tends to overrepresent users with higher end rigs. If you look at the non-gaming market, the hardware level tends to be considerably worse, and as a result a lot of productivity programs still ship plain SSE2 as their baseline required target.
> Not sure why you keep referring to Skylake since Intel hasn't shipped a Skylake-based CPU in years, and the most problematic CPUs currently shipping that lack AVX-512 support are several architectural generations ahead of Skylake, for both the P-cores and E-cores.
I could have swore Intel's first E-cores were skylake based. My bad. Point is that its Intel issue, not AVX-512 issue.
> They still ship huge volumes of CPUs, especially in laptops where AMD is still underrepresented.
They physically shipped them, yes, but the products were lackluster. I don't recall when Intel had a good desktop CPU last time.
> Targeting a discrete GPU is different than not supporting integrated GPUs at all, which would be analogous.
The analogous case would be targeting a specific graphics API feature level, say, D3D 9.x, because that's what the majority of GPUs on the market support. Except game developers somehow figured out that they can support multiple D3D feature levels instead of being permanently stuck targeting whatever the majority happens to have.
Why not? To save die space?
> However, if you are distributing the binaries for other people to run, that’s not really an option.
This all depends on what kind of software you're making. A lot of games set their requirements about 5 generations back, like FC 27 where the minimum is a Ryzen 1600. That lets them use AVX2 unconditionally and prevent complaints from users who tried to run it with a super old CPU.
Then you get whole Linux distros like CachyOS and Clear (RIP) that rebuild the world for each architecture level and have them as separate variants. I think it still counts as binaries for other people.
Also the state of SIMD in Cranelift is also very WIP. They pretty much just support a subset of 128bit vectors with some rare exceptions.
The question for me is whether portable simd will result in faster code than plain auto-vectorisation; for the simplest loops auto has me beat (the few times I've tried it), but I imagine as the complexity grows I'll be more likely to try do something that breaks auto-vectorisation, and it'll be more obvious to me when I do that in portable simd.
Warning, Nitpick. Saying "hardware [...] is cheap" and "instruction decoding is expensive" (implied) is a minor contradiction. You can actually duplicate instruction decoding just fine, it's the thing before it where things go to hell: Instruction fetching.
Nothing prevents you from building a computer that can fetch, decode and execute 16 instructions at once, assuming they don't all write to the same register.
But if you want to do the same add repeated 16 times you'll need 16 times more program memory and 16 wider read ports on your caches and so on. SRAM is really expensive so this strategy will waste a lot of area on memory that you probably didn't need in the first place. I say this as someone who had to design a chip in university and basically you couldn't even find the primitive CPU in-between the massive SRAM blocks. By reusing the same instruction you can now increase your compute to memory ratio in terms of area.
Just a heads up for people who want to know why SIMD is a thing. I'm not criticizing the article, I just want more people to realize the pain that SRAM represents to chip designers.
Just for the sake of curiosity it would be nice to have a peek at what SIMD in Rust looks like on RISC-V today. Yes, even if it requires some "obscure compiler flags" for now (while we wait for the Oilsm extension).
Downside: It's currently x86 only.
It's the closest thing to Google's Highway.
(I am not connected to Fearless SIMD in any way other than as a user.)
You can either have performance (=write manual ASM for each platform), or portability, but not both.
What so-called "portable SIMD" libraries give you is "portable auto-vectorization". "Portable performance" is a global property of the algorithm. Relying on auto-vectorization will result in e.g. sub-optimal register spills in practice. The microbenchmarks will look great, though. ;)
But there's another point in the tradeoff space. One of the explicit design decisions in Fearless SIMD is to support "downcasting," or specialization to a specific microarchitecture. At least for the kind of problems I've worked on, even when you're doing something fancy with arch-specific permutations or what not, the majority of the operations will be pretty vanilla, and can be expressed well in the portable subset.
So you can think of a library like Fearless SIMD as enabling your extreme optimization use case, just more ergonomically.
Of course, this depends on LLVM compiling intrinsics to assembly efficiently. That hasn't always been the case, and is not perfect now (a number of issues have been filed against rustc and LLVM while developing Fearless SIMD), but is pretty good.
As always, though, you do have to measure performance, and I frequently look at the assembler output to double-check that it's doing the right thing. The day of "fire and forget" portable SIMD has not yet arrived.
For example, consider you want to compute the average of a list of floats. The compiler cannot autovectorize this, because float addition is not commutative. However, it's much faster to do component-wise addition in groups, then a horizontal sum at the end, and then divide. Whether it matters depends on your use case, and the compiler unfortunately can't read your mind, so it has to be conservative.
Gcc and llvm can tell you if they can’t Auto vec a function, maybe rust could turn this into an error at comptime.
Nowadays you can even get AI to write intristics and it works just fine, the portable libraries/autovec aren't really a serious player here.
Portability is also overstated - see the recent shift where Spotify decided to make native Android/iOS apps again instead of React Native. Usually, the number of relevant platforms is somewhere between 2 and 3, so portability concerns are more theoretical than real.
I'm getting 50x faster code with manual ASM. That's the difference between audio code that runs in realtime and code that does not.
The competing implementations use SIMD and native code, autovec works nicely there. I symbolically invert the LinAlg system at compile time.
IMO this is just people over-indexing on 10 year old GCC. Modern LLVM versions (and even GCC) mostly do good things out of the box. The hard part for the compiler is the vectorization strategy, so using portable intrinsics gives the compiler the shape and it generally does a very good job from there.
Thoughts on an abstraction over ARM and x86, at 128, 256, and 512-bit widths which, either in a manual or automatic way (The latter more challenging) makes your floating point computations 4-16x faster with minimal restructuring? I think that's doable, and a nice goal of SIMD.
Except in languages with a JIT compiler
Granted, the number of cases this distinction matters is relatively small, making a function faster only makes a program appreciably faster if that function is a bottleneck.
Erm, not sure how direct you want. But at least in dotnet you can use the Vector64, Vector128, Vector256, and Vector512 types [1] where each method gets effectively directly converted to a raw SIMD instruction. So, Vector512.LoadAligned call will be replaced with a raw SIMD register load instruction - supported by the CPU it is running on; and generally a JIT compiler will spot common patterns-of-use to optimise those too. There's no runtime check to see what is supported and no per-function-branching. It's as close to the CPU as you can get really (in a compiled language). Maybe I'm missing something?
If you just mean the difference between hand-coded assembly and the output of an optimising compiler, then sure, you can always be better with hand-coded assembly.
[1] https://learn.microsoft.com/en-us/dotnet/api/system.runtime....
[1] https://learn.microsoft.com/en-us/dotnet/api/system.runtime....
Azul and OpenJ9 additionally have server JITs, which widen the abilities of optimisations are available.
Additionally the ART cousin also does its own thing.
Each JVM implementation has its own approach how to do auto-vectorisation or mark intrinsic methods.
It is no different than talking about Ada, Fortran, COBOL, C, C++ and co compilers versus what ISO defines in the language standard.
But yeah to be fair if you are at that point, you probably want to go fully non-portable anyway. Especially with AI.
Has anyone even figured out how to do vector stuff (SVE/RVV) without assembly?
There's also Halide, where you write the algo but the framework gets you the scheduling and SIMD.
https://val.markovic.io/articles/philbin-the-safest-and-fast...
Performance is competitive with Zig, without using the "unsafe" keyword or nightly Rust.
Nice.
NGL, SIMD in Rust has not been a great experience. Philbin uses the "CPU Feature Tokens" pattern that the Fearless SIMD team (which the root article's author is a part of) came up with: https://shnatsel.medium.com/safe-simd-in-rust-even-on-the-in...
The pattern provides a safe abstraction for runtime CPU detection and SIMD dispatch. It's gnarly AF, but Fearless SIMD 1.0 came out a few days ago so people won't have to implement it themselves anymore.
I'd recommend grabbing that lib and giving SIMD-in-Rust another shot.