> The only case I can currently fore see where using LMUL=1 and manually unrolling instead will likely be always beneficial is vrgather operations that don't need to cross between registers in a register group (e.g. byte swapping).
This is somewhat of a major problem with some kernels. RVV is quite spartan in the permute options it offers, often forcing you to use vrgather for many things. And as you suggest, vrgather doesn't scale well, so sticking to LMUL=1 seems sensible in a lot of cases. (this will also be a problem for RISC-V implementations that aim for longer vectors)
Honestly LMUL>1 could be more useful if more permute instructions were offered, particularly a restricted shuffle (like VPSHUFB on AVX or TBLQ on SVE2.1).
Stuff like vector constants can be more costly with LMUL>1, since they must consume more than one register.
On big cores, the only benefit LMUL>1 gives (other than design concepts like how widen work, or shuffling across vectors etc) is a code size reduction. Which is quite a dubious benefit for a somewhat complex feature.
Maybe smaller cores can extract more out of it, but it's not something I'm too knowledgeable about.