Every system I'm aware of has 128-bit SIMD implemented: either SSE (x86), NEON (ARM), or AltiVec (POWERPC). As such, 128-bit SIMD is the "reliably portable" SIMD operation.
Of course, for fastest speeds, you need to go to the largest SIMD-register size for your platform: 512-bit for some Intel processors, rumored AMD Zen4 and Centaur chips. 256-bit for most Intel / AMD Zen3 chips. 128-bit for Apple ARMs, 512-bit for Fujitsu A64 ARMs, etc. etc.
> And these are the fastest kinds of instructions :-)
And it should be noted that modern SIMD instructions execute at one-per-clock-tick. So the 64-bit instructions certainly are fast, but the SIMD instructions tie if you're using simple XOR or comparison operations.
--------
This style of code might even be "reliably auto-vectorizable" on GCC and CLANG actually. I wonder if I could get portable auto-vectorizing C from these examples.