As to libraries: I'm the main author of github.com/google/highway which provides 'portable intrinsics' so you only have to write your code once for all platforms.
It includes ready to use sorting and searching algorithms in hwy/contrib.
wow this looks very interesting, I noticed the release passed 1.0 milestone, how did you unify all those intrinsics? I'm particularly interested in RVV1.0.
With a lot of the JPEG XL code written, when SVE and RVV were being introduced/discussed, I realized that pretty much the same code would work there, too: we just needed to replace a constexpr VecClass::kLanes with Lanes(), which is what Highway now does.
Beyond the initial set of efficient-everywhere ops, we've added some (such as AbsDiff) that help on say Arm without hurting other platforms, nor can user code do better itself. There are also a few (ReorderWidenMulAccumulate) that define a non-obvious relaxation of the interface which is more efficient for all platforms than holding to any one platform's interface.
For RVV, one remaining concern is the VSETVLI - the revised intrinsics now require an avl argument, which Highway creates in each function. It's not yet clear whether the compiler will be smart enough to (peephole-?)optimize away the duplicates.
1: https://github.com/google/highway/blob/master/g3doc/instruct...
It's a lot like the difference between writing C or assembly. Even today, sometimes you need asm to make the code go real quick because what the compiler spits out isn't always optimal (but for the general case, it's quite good).
https://www.tomshardware.com/news/intel-nukes-alder-lake-avx...
Wide-datapath designs will generally have some tradeoffs like needing to ramp up power delivery, so there will be things like initial instruction latency for wide-datapath instructions, etc. That's kind of inevitable; I suspect the same will be true of the new fancy "variable length" vector ISAs if the underlying implementation and vector usage is wide enough, too.
Also: you don't need to use wide 512-bit vectors with AVX-512! You can use the instruction set with 128-or-256/bit vectors just fine.
https://travisdowns.github.io/blog/2020/08/19/icl-avx512-fre...
Actually supposedly Skylake-X/Xeon-W had much lower AVX downclocking than Skylake-SP too... InstLat64 made a tweet at one point showing this was 10-20% for workstation vs 30-40% for server iirc. Tweet has been removed unfortunately.
Intel really throttled it down on server chips, for whatever reason. Probably didn't want datacenter chips to run the Unlimited Voltage that was necessary for full-clock dual-unit AVX-512 on 14nm.
[ General observation, not directed at parent comment: ]
Frequency throttling, even on the most affected Skylakes, has always been a non-issue if you run say 1ms worth of continuous SIMD instructions. How could a 10-40% drop negate speedups from 2x vector width plus double the registers and a much more capable instruction set?
It is time we buried this myth :)
My 10980XE runs AVX2 at 4.2 GHz all core, and AVX512 at 4GHz.