I'd think that a fundamental problem is memory I/O bandwidth, and (depending on cache policy) poor cache hit ratios.
I'd think that a fundamental problem is memory I/O bandwidth, and (depending on cache policy) poor cache hit ratios.
A typical example would be: you have an array of objects of constant size, and you’re reading a double field from a constant offset within each object. The hardware prefetcher will “recognize” this access pattern and prefetch that offset every sizeof(obj) bytes.
The major downsides (vs. a struct-of-arrays design with full spatial locality) are:
1. Every prefetch pulls a full cache line, but the cache line will include data you don’t need. In this example, every cache line might have 64 bytes of data, but you only needed the one double field (8 bytes) - the rest is not useful. If you were iterating over an array of doubles you could have pulled 8 double fields in a single cache line.
2. Specific performance is hardware-dependent, so it’s hard to guarantee performance on e.g. low-end cores, short loops, or unusually long strides.
You seem to be assuming that a vectorized loop over an array of structs only ever uses one element from the struct and ignores the rest. In that case, then yes, you will get some bandwidth problems, but modern hardware prefetchers are so good that a vectorized loop that only looks at each Nth byte or whatever is actually much faster than one might think. I'll also note that the serial (non-SIMD) case also suffers here, but often not quite so badly as the SIMD case suffers.
However, there's lots of cases where you want the whole struct. E.g. complex numbers are typically a struct of two real numbers, and SIMD arithmetic on vectors of structs of complex numbers is able to work just fine with the full theoretical SIMD performance improvement.
That is the worst case scenario of my concern with memory bandwidth.
More generally, I'd imagine that any good roofline analysis would need to consider this data-layout issue.
> but modern hardware prefetchers are so good that a vectorized loop that only looks at each Nth byte or whatever is actually much faster than one might think.
I believe you. I wasn't really thinking about how efficiently the processor can pull the correct bytes into the register lanes. I was just focusing on the fundamental memory <--> CPU xfer part of the process.
TL;DR: I'm trying to get better at using roofline analysis in my optimization efforts.
Is that not the case?