By that, do you mean the AVX2 'gather' type instructions? If not, I'd be interested to know what those techniques are.
As for AVX2 gathers, I had to look this up recently and it sounds like they're about as fast as manually unpacking the vector and performing scalar loads. That is to say, they're decidedly not fast. On the other hand, it sounds like (as of Skylake) they're bottlenecked on accesses to the L1 cache, so they're about as fast as they reasonably could be.
Source: https://stackoverflow.com/questions/21774454/how-are-the-gat...
Not sure about performance on Zen, but I would imagine it's similar?