SIMD Everywhere Optimization from ARM Neon to RISC-V Vector Extensions
arxiv.org
arxiv.org
Direct link to PDF: https://arxiv.org/pdf/2309.16509.pdf
https://doc.rust-lang.org/std/simd/struct.Simd.html
Libraries implemented in languages without these semantics will greatly benefit from this.
https://en.cppreference.com/w/cpp/experimental/simd
This proposal has been around for a while a but it recently got some new momentum and seems to be on track for c++26. Gcc ships a version today for those wanting to try it.
It's difficult to understand the value proposition of std::experimental::simd. Standardization has taken many years, lost the "load/store" function naming which just about everyone everywhere(?) is using, and only resulted in ~30 ops [1] that are also mostly achievable with compiler builtins or perhaps even autovectorization.
Four years ago, we stepped in with Highway to fill the gap; it now has about 250 ops [2] and an active community. I have not wanted to step on any toes, but wonder whether the ship has by now sailed on this?
[1]: https://en.cppreference.com/w/cpp/header/experimental/simd [2]: https://github.com/google/highway/blob/master/g3doc/quick_re...
Not sure where the momentum is coming from. The proposal has a new maintainer now, new drafts have been released, and there's a stated goal to get this included in C++26.
For me, the appeal for experimental/simd would be it being part of the stdlib. Have you considered proposing Highway for stdlib inclusion? From your description it appears to be a more complete implementation. I realize it's a long and arduous process, but I think it's the only way to make something truly the default implementation.
(And thanks for the link. I was not aware that Highway existed.)
At the time we open-sourced Highway, the standardization process had already started and there were some discussions.
I'm curious why stdlib is the only path you see to default? Compare the activity level of https://github.com/VcDevel/std-simd vs https://github.com/google/highway. As to open-source usage, after years of std::experimental, I see <200 search hits [1], vs >400 for Highway [2], even after excluding several library users.
But that aside, I'm not convinced standardization is the best path for a SIMD library. We and external users extend Highway on a weekly basis as new use cases arise. What if we deferred those changes to 3-monthly meetings, or had to wait for one meeting per WD, CD, (FCD), DIS, (FDIS) stage before it's standardized? Standardization seems more useful for rarely-changing things.
1: https://sourcegraph.com/search?q=context:global+std::experim...
2: https://sourcegraph.com/search?q=context:global+HWY_NAMESPAC...
There is Highway (https://github.com/google/highway) however, that does support dynamically-sized SIMD.
The point is to not have to write repetitive source code many times.
And yes, it's not particularly large of a cost, other than it being an extremely pointless waste of space given that it is possible to have just one variant that covers them all.
Though, it would become significantly more problematic if you wanted to target different extension groups too (which you would quite likely want to some extent) as those'd multiply with all the length targets - SVE vs SVE2 vs more future extensions, and on RVV there's just a lot (Zvfh & Zvfhmin for FP16, Zvbb for extra bitmanip stuff, many more here[1]; and potentially at some point there could be an extension that uses a wider encoding scheme to inline vsetvl fields & allow masking by registers other than v0, which could benefit everything)
[1]: https://github.com/riscv/riscv-crypto/blob/c8ddeb7e64a3444dd...
Compiled languages like Rust, C++ and Zig cannot detect the hardware because they have no runtime right? Could a language like Go add the simd semantics and detect the support vector size?
The problem is that a "Simd<i32, 4>" will always have 4 elements, but you'd need a "Simd<i32, whatever the hardware has>" type, which has significant impact on what is possible to do with such a type.
So technically it sounds feasible but all of the languages like Zig, C++ and Rust picked a simpler approach. Is it simply a first step to a more abstract approach?
And some things you just can't really "generalize" to scalable vectors. e.g. you can store Simd<i32,4> in a struct or global variables, or initialize with, say, [3,2,1,0], but none of those things are possible with scalable vectors (globals/struct fields need a known size, and initializing with a hard-coded list of elements doesn't make much sense if you don't even know how many elements you'll need).
Zig intends to support a similar feature but doesn't yet, at least not built into the language (you could certainly express this if you tried hard enough). I don't know about Rust, but I would be very surprised if it can't do this.
edit: I think I replied to the wrong comment >.<
For example:
However, I'm pretty sure OpenCV has their "universal intrinsics" and RISC-V with scalable vector registers is supported in the latest OpenCV version
Universal intrinsics (docs not updated): https://docs.opencv.org/4.x/d6/dd1/tutorial_univ_intrin.html Scalable RVV support: https://github.com/opencv/opencv/pull/22179
I hope they have real hardware performance numbers for the rv summit talk.
There have been many SIMD abstraction layers created in the past but none of them will beat the raw speed of handwritten assembly. Try and implement something like vpternlogd in one of these abstraction layers.
E.g. in asm you can run the same instruction sequence with different vtype (element width and LMUL).
It's an extremely fun idea (primarily just for code size though), but thankfully (?) its usability is restricted by load/store instrs hard-coding the element type, so the main use of this would end up for switching LMUL, which has very limited usefulness.
What Highway can support is generating multiple loops of different vtype from the same code, which effectively achieves the same thing, at the cost of machine code duplication.
I currently have a quite usefull use case for it, I'm concerting utf8 to utf32 and if I've got an average utf8 character size of above 2 I could reduce the LMUL for that loop iteration.
This shouldn't actually improve performance that much in good rvv implementations, since you can use vl and not LMUL to schedule your execution units. Sadly this is currently not the standard, and ara is the only implementation, that does this I know of.
I think this wouldn't even be about code size reduction, consider an input, where there is basically a 50/50 probability LMUL can be reduced, that would be horrible for the branch predictor, but with only a branch over vsetvl, this could behave as a conditional vsetvl via instruction fusion. We'll have to see if such optimization become relevant once there is more hardware out there.
Also instruction scheduling. Low-end Cortex will probably be in-order till the end of time...
In general, kierank is right: if you want to full optimize something, and you know what you want the actual code to look like, just write it in assembly. Nothing else gives you full control over loads and stores, and anything else leaves you at the mercy of some future compiler "optimization" stepping in to defeat you.
Interesting. I thought you'd be at the mercy of the instruction scheduler for aggressive OoO cores.
Despite all the other comments, this doesn't appear to be intended to be used to write SIMD across multiple platforms? Rather, it's to quickly port codebases with lots of existing platform-specific intrinsics to a new platform? For this paper in particular, so that RISC-V can run somewhat optimized code without having to spend thousands of man-years writing new RVV code.
BTW this reminds me of a colleague grumbling that what should have been a 20-minute patch to ffmpeg took a day, because it was written in assembly.
It is also quite possible to have large slowdowns due to assembly - all it takes is to forget a v prefix (VEX encoding), whereas intrinsics take care of that.
The lightweight macro layer in ffmpeg takes care of v prefixes.
In FFmpeg, x264 and dav1d there are many different examples of code that couldn't be written in intrinsics or other abstraction layer.
https://twitter.com/FFmpeg/status/1705543447245988245?t=Ul9e...
hm, I vaguely remember there was a vzeroupper problem, perhaps one fell through the cracks.
Interesting, can you share more details on the magic? Looks mainly like function call overhead. If functions aren't called often, we can inline (by moving into headers or enabling LTCG/LTO).
If they are called often, are visible to the compiler, have internal linkage, but shouldn't be inlined, I'd be curious to learn why, and also why the compiler is then generating the full prolog/epilog.
I myself implemented one in the SSE4/Altivec days (later extended to AVX, AVX512 and NEON). There were only a few options then, but now everyone seems to be doing it.
There should be a simd.h header in the C standard library that contains typedefs for vector types, and various functions to operate on them as well as Operators for them.
Like my _Operator <symbol> <function name>; proposal, which requires no mangling.