Grabbing the sources from https://www.nayuki.io/page/fast-sha2-hashes-in-x86-assembly and compiling them more or less as recommended (and unlike they are compiled in the talk):
$ clang -O3 sha256-test.c sha256.c -o sha256-test ; for i in 1 2 3; do ./sha256-test ; done
Self-check passed
Speed: 197.2 MB/s
Self-check passed
Speed: 196.1 MB/s
Self-check passed
Speed: 196.1 MB/s
$ gcc -O3 sha256-test.c sha256.c -o sha256-test ; for i in 1 2 3; do ./sha256-test ; done
Self-check passed
Speed: 209.7 MB/s
Self-check passed
Speed: 209.2 MB/s
Self-check passed
Speed: 209.1 MB/s
This is a baseline for us. It has nothing to do with Zig and nothing to do with Andrew's machine (hardware or compiler versions).But wait, the page above suggests that adding -march=native might help. Indeed it does:
$ clang -O3 -march=native sha256-test.c sha256.c -o sha256-test ; for i in 1 2 3; do ./sha256-test ; done
Self-check passed
Speed: 255.6 MB/s
Self-check passed
Speed: 259.0 MB/s
Self-check passed
Speed: 254.3 MB/s
$ gcc -O3 -march=native sha256-test.c sha256.c -o sha256-test ; for i in 1 2 3; do ./sha256-test ; done
Self-check passed
Speed: 275.1 MB/s
Self-check passed
Speed: 268.4 MB/s
Self-check passed
Speed: 270.0 MB/s
In the talk Andrew suggests that the difference might be due to using rorx instructions, which Zig might be able to do due to aggressive loop unrolling. Does -funroll-all-loops help GCC? It turns out that it doesn't, and that it cannot, on this program, because the C code for sha256_compress is already fully unrolled.But anyway, are we using rorx instructions at all? We are, but only with -march=native:
$ clang -O3 -S sha256.c -o - | grep -c rorx
0
$ clang -O3 -march=native -S sha256.c -o - | grep -c rorx
542
And: $ gcc -O3 -S sha256.c -o - | grep -c rorx
0
$ gcc -O3 -march=native -S sha256.c -o - | grep -c rorx
576
[Edit: Changed the grep from "ror" to "rorx", which changed GCC's numbers a bit; it does generate "ror" without x without -march=native.]So. Testable theories:
(a) On Andrew's machine, the same setup but with Clang using -march=native would outperform or at least match Zig.
(b) Zig's compiler internally uses the equivalent of -march=native, possibly implicitly, at least in --release-fast mode.
Nothing here is meant to imply that Andrew is dishonest. There are just lots of variables to take into account, and sometimes we don't. Also, "slower/faster than C" by a few percent is not very meaningful if even C compilers disagree by 6% or so, and the same C compiler with the right flags disagrees with itself by a lot more.