Here’s how I would do that in AVX2: https://gist.github.com/Const-me/eed10bfe690b5804d2fc8266e02...
I wonder how does the performance compare to your version.
Here’s how I would do that in AVX2: https://gist.github.com/Const-me/eed10bfe690b5804d2fc8266e02...
I wonder how does the performance compare to your version.
Using srand(0) as in your gist; changing int8_t to char ; using int for sum:
BM_hn 10155300 ns
BM_avx512 9218652 ns
-----
BM_avx512 9372319 ns
BM_hn 10428792 ns
Your code is getting ~18-19GB/s read bandwidth compared to roughly ~21GB/s for my code.
I wonder how much of that is due to interleaved summing vs not compared to AVX512 vs AVX2.
Assuming none of us screwed up too badly, that gives us about 10% profit from AVX512 compared to AVX2. Not sure that’s a good enough win to justify vendor lock-in to Intel.
P.S. int8_t is a typedef for char, that change was stylistic and should not affect performance whatsoever. I tend to avoid char data type for numbers, as opposed to characters.
Update: you can try to make another AVX512 version with _mm512_sad_epu8 like I did for AVX2, instead of that integer dot product. I won’t be surprised to find out vpsadbw is faster than vpdpbusd, fundamentally addition is simpler than multiplication.