RISC-V Vector Extension for Integer Workloads: An Informal Gap Analysis
gist.github.com
gist.github.com
> This should be mostly mitigated because the base V extension only has the full lane-crossing-gather, hence a lot of software will use lane-crossing operations and force vendors to implement it competitively.
Is that actually what's happening? From the discussion you link (https://github.com/llvm/llvm-project/pull/104574), it seems like the X280 behavior is currently still widespread. Do we know when the competitive perf might materialize?
The linked discussion was about LMUL>1, which doesn't matter because you can unroll it into multiple LMUL=1 vrgathers if you don't cross lanes. That even uses fewer vector registers. LMUL>1 vrgather should imo only be used if you need to actually cross vector register lanes.
I think a fast 128-bit lane shuffle is certainly unnegotiable for application class processors, but I'm not sure a dedicated instruction is needed for them, because for them a fast LMUL=1 vrgather is similarly important. You won't see many processors with larger then 512 VLEN running regular software from the binary ecosystem and if they do, they'd better execute a LMUL=1 shuffle quickly, especially if it's within 128/256/512-bit lanes, whatever your internal execution unit width is.