A curious SIMD assembly challenge: the zigzag
x264dev.multimedia.cx
x264dev.multimedia.cx
I'm a n00b wrt SIMD, but is it really better to do unaligned loads instead of aligned loads + shift or shuffle?
According to latency tables on Core 2 loading 128 bits from memory to an SSE register
* movdqa: latency of 3cy, reciprocal latency 1cy
* movdqu: latency of 3-8cy, reciprocal latency 4cy
Also movdqu seems to have 9 µops vs. only 1 µops for movdqa. Wouldn't there be bad throughput on those chunks of 4 movdqu?[All this only according to Intel Manuals and Agner's tables, aka no testing]
How do you guys test implementation performance?
Thanks!
It may be worthwhile for Core 2 to do a bunch of aligned loads and palignr them together, but I didn't feel like testing, as it would certainly have been slower on my i7. Patches welcome ;)