Last time i looked at intel scatter/gather I got the impression it only works for a very narrow use case, and getting it to perform wasn’t easy. Did I miss something?
I would say that for SIMD the situation is basically the same. gather/scatter don't magically make the memory hierarchy a non-issue, but they're no longer adding any unnecessary pain on top.
> The VGATHER instructions are implemented as micro-coded flow. Latency is ~50 cycles.
https://www.intel.com/content/www/us/en/content-details/8141...