There have been many SIMD abstraction layers created in the past but none of them will beat the raw speed of handwritten assembly. Try and implement something like vpternlogd in one of these abstraction layers.
There have been many SIMD abstraction layers created in the past but none of them will beat the raw speed of handwritten assembly. Try and implement something like vpternlogd in one of these abstraction layers.
BTW this reminds me of a colleague grumbling that what should have been a 20-minute patch to ffmpeg took a day, because it was written in assembly.
It is also quite possible to have large slowdowns due to assembly - all it takes is to forget a v prefix (VEX encoding), whereas intrinsics take care of that.
The lightweight macro layer in ffmpeg takes care of v prefixes.
In FFmpeg, x264 and dav1d there are many different examples of code that couldn't be written in intrinsics or other abstraction layer.
https://twitter.com/FFmpeg/status/1705543447245988245?t=Ul9e...
hm, I vaguely remember there was a vzeroupper problem, perhaps one fell through the cracks.
Interesting, can you share more details on the magic? Looks mainly like function call overhead. If functions aren't called often, we can inline (by moving into headers or enabling LTCG/LTO).
If they are called often, are visible to the compiler, have internal linkage, but shouldn't be inlined, I'd be curious to learn why, and also why the compiler is then generating the full prolog/epilog.
In general, kierank is right: if you want to full optimize something, and you know what you want the actual code to look like, just write it in assembly. Nothing else gives you full control over loads and stores, and anything else leaves you at the mercy of some future compiler "optimization" stepping in to defeat you.
Interesting. I thought you'd be at the mercy of the instruction scheduler for aggressive OoO cores.
Also instruction scheduling. Low-end Cortex will probably be in-order till the end of time...
E.g. in asm you can run the same instruction sequence with different vtype (element width and LMUL).
It's an extremely fun idea (primarily just for code size though), but thankfully (?) its usability is restricted by load/store instrs hard-coding the element type, so the main use of this would end up for switching LMUL, which has very limited usefulness.
What Highway can support is generating multiple loops of different vtype from the same code, which effectively achieves the same thing, at the cost of machine code duplication.
I currently have a quite usefull use case for it, I'm concerting utf8 to utf32 and if I've got an average utf8 character size of above 2 I could reduce the LMUL for that loop iteration.
This shouldn't actually improve performance that much in good rvv implementations, since you can use vl and not LMUL to schedule your execution units. Sadly this is currently not the standard, and ara is the only implementation, that does this I know of.
I think this wouldn't even be about code size reduction, consider an input, where there is basically a 50/50 probability LMUL can be reduced, that would be horrible for the branch predictor, but with only a branch over vsetvl, this could behave as a conditional vsetvl via instruction fusion. We'll have to see if such optimization become relevant once there is more hardware out there.
Despite all the other comments, this doesn't appear to be intended to be used to write SIMD across multiple platforms? Rather, it's to quickly port codebases with lots of existing platform-specific intrinsics to a new platform? For this paper in particular, so that RISC-V can run somewhat optimized code without having to spend thousands of man-years writing new RVV code.