Are there cases where the benchmark doesn't hold true? Different architectures?
I'm inclined to use this on some less security sensitive things I need a hash for.
Good for files and similar if your data stream is long enough to effectively use the tree structure.
Yes, that 5x figure represents a big speedup that comes from using SIMD (specifically AVX-512) to compress multiple blocks in parallel. To take full advantage of AVX-512, BLAKE3 needs to have at least 16 KiB of input. So for any input shorter than that, the speedup compared to BLAKE2b/s will be smaller. For inputs less than 2 KiB, the only speedup is the ~1.4x that you get from the round reduction.
Architectures other than x86 tend to have less in the way of SIMD. ARM has NEON, but that's currently only 128 bits wide. So the big single-threaded speedups are only on x86 today. (For multithreading speedups, the architecture doesn't matter as much. Just the number of cores you have.)
> I'm inclined to use this on some less security sensitive things I need a hash for.
This makes sense insofar as BLAKE3 is a new design, and it pays for cryptographic applications to be conservative. But to be clear, if BLAKE3 turns out not to uphold the same security properties as SHA-2 or BLAKE2s, that would represent a catastrophic failure of the design. (It's possible that some flaw could be discovered that affects BLAKE2 and BLAKE3 equally, but that the lower round count of BLAKE3 would make the flaw more severe. In that case, I'd expect everyone would also want to migrate off BLAKE2 pretty quickly anyway.)
Makes me wonder if BLAKE3 will be changed to suit, or if Intel chips are just going to suck when it come to performance with the standard?
The story is different with SHA-2, which is entirely serial by design. There's no possibility of parallelism between different blocks of a single SHA-2 input. That means that there's little incentive to create a parallelized hardware implementation of SHA-2, and I'm not aware of any. In contrast, BLAKE3 is more like AES-CTR and ChaCha in that (given a long enough input) it can take full advantage of any degree of parallelism. For this reason, given reasonably modern SIMD support like AVX-2, BLAKE3 tends to be faster than hardware-accelerated SHA-256, even before multithreading gets into the picture.
No. Most Intel CPUs lack SHA-NI extensions. It looks like they may have finally started adding it to their 2020+ mainstream CPUs, but for a long time it was only present on their low-end Atom and similar CPUs. AMD CPUs have all had SHA-NI support since Zen.