I'd think that a fundamental problem is memory I/O bandwidth, and (depending on cache policy) poor cache hit ratios.
Is that not the case?
A typical example would be: you have an array of objects of constant size, and you’re reading a double field from a constant offset within each object. The hardware prefetcher will “recognize” this access pattern and prefetch that offset every sizeof(obj) bytes.
The major downsides (vs. a struct-of-arrays design with full spatial locality) are:
1. Every prefetch pulls a full cache line, but the cache line will include data you don’t need. In this example, every cache line might have 64 bytes of data, but you only needed the one double field (8 bytes) - the rest is not useful. If you were iterating over an array of doubles you could have pulled 8 double fields in a single cache line.
2. Specific performance is hardware-dependent, so it’s hard to guarantee performance on e.g. low-end cores, short loops, or unusually long strides.
You seem to be assuming that a vectorized loop over an array of structs only ever uses one element from the struct and ignores the rest. In that case, then yes, you will get some bandwidth problems, but modern hardware prefetchers are so good that a vectorized loop that only looks at each Nth byte or whatever is actually much faster than one might think. I'll also note that the serial (non-SIMD) case also suffers here, but often not quite so badly as the SIMD case suffers.
However, there's lots of cases where you want the whole struct. E.g. complex numbers are typically a struct of two real numbers, and SIMD arithmetic on vectors of structs of complex numbers is able to work just fine with the full theoretical SIMD performance improvement.
That is the worst case scenario of my concern with memory bandwidth.
More generally, I'd imagine that any good roofline analysis would need to consider this data-layout issue.
> but modern hardware prefetchers are so good that a vectorized loop that only looks at each Nth byte or whatever is actually much faster than one might think.
I believe you. I wasn't really thinking about how efficiently the processor can pull the correct bytes into the register lanes. I was just focusing on the fundamental memory <--> CPU xfer part of the process.
TL;DR: I'm trying to get better at using roofline analysis in my optimization efforts.
https://documentation-service.arm.com/static/6530e5163f12c06... (PDF)
Last time I tried using gathers on AVX2, performance was comparable to doing scalar loads.
Gathers on AVX2 used to be problematic, but assume it shouldn't be the case today especially if the lane-crossing is minimal? (if you do know, please share!)
When using SIMD you must either use SoA or AoSoA for optimal performance. You can sometimes use AoS if you have a special hand coded swizzle loader for the format.
You can write them yourself also but I'd add a verifier that checks the output with scalar code as it can be tricky to get correct.
Intel had an article up for the 3x8 transpose, but it seems to no longer exist so i'll just post the psuedo code
//xyz -> xxx
void swizzle3_AoS_to_SoA(v8float &x, v8float &y, v8float &z) {
v8float m14 = interleave_low_high<1, 2>(x, z); //swap low/high 128 bits
v8float m03 = blend<0, 0, 0, 0, 1, 1, 1, 1>(x, y); //_mm256_blend_ps 1 cycle
v8float m25 = blend<0, 0, 0, 0, 1, 1, 1, 1>(y, z);
//shuffles are all 1 cycle
__m256 xy = _mm256_shuffle_ps(m14, m25, _MM_SHUFFLE(2, 1, 3, 2)); // upper x's and y's
__m256 yz = _mm256_shuffle_ps(m03, m14, _MM_SHUFFLE(1, 0, 2, 1)); // lower y's and z's
v8float xo = _mm256_shuffle_ps(m03, xy, _MM_SHUFFLE(2, 0, 3, 0));
v8float yo = _mm256_shuffle_ps(yz, xy, _MM_SHUFFLE(3, 1, 2, 0));
v8float zo = _mm256_shuffle_ps(yz, m25, _MM_SHUFFLE(3, 0, 3, 1));
x = xo;
y = yo;
z = zo;
}