I've tested XXH3 using xxhash's built-in benchmark tool with clang-7.0.1 and gcc-8.2.1 on an Intel i9-9900K. The processor was otherwise idle, and was running at 5 GHz. The command I tested is `xxhsum -b5i10`.
SSE2: CFLAGS=-O3
AVX2: CFLAGS="-O3 -mavx2"
ARCH: CFLAGS="-O3 -march=native"
Compiler Mode Speed
gcc-8 SSE2 32.8 GB/s
clang-7 SSE2 36.5 GB/s
gcc-8 AVX2 44.1 GB/s
clang-7 AVX2 68.3 GB/s
gcc-8 ARCH 60.9 GB/s
clang-7 ARCH 69.7 GB/s
gcc nearly catches up with clang when compiled with -march=native, but with only -mavx2 it isn't performing as well.