Using srand(0) as in your gist; changing int8_t to char ; using int for sum:
BM_hn 10155300 ns
BM_avx512 9218652 ns
-----
BM_avx512 9372319 ns
BM_hn 10428792 ns
Your code is getting ~18-19GB/s read bandwidth compared to roughly ~21GB/s for my code.
I wonder how much of that is due to interleaved summing vs not compared to AVX512 vs AVX2.