A trick not mentioned here: for Python devs, import “orjson” instead of the json standard library; it is usually a drop-in replacement.
A trick not mentioned here: for Python devs, import “orjson” instead of the json standard library; it is usually a drop-in replacement.
They may be slow only in programs which do not check for hardware support and which do not use the dedicated hardware instructions (which is the case in many programs, because Intel Skylake derivatives have been the most popular CPUs during many years, and they were the only modern CPUs without hardware support for SHA-1/SHA2-256, so most developers did not bother to optimize their programs for the other CPUs).
OpenSSL is one of the few programs which use the hardware instructions when available. This makes, e.g., "openssl dgst -sha256" much faster than "sha256sum" on recent CPUs.
When using the hardware instructions, SHA-1 and SHA2-256 are faster than many non-cryptographic hashes and only 1 cryptographic hash is faster: BLAKE3.
However, it must be kept in mind that BLAKE3 is much faster than any other cryptographic hash only because it distributes the computation on all CPU cores. So the much higher speed is accompanied by a much higher CPU utilization.
If you have something else that should be done in parallel with the hash computation, using BLAKE3 does not necessarily reduce the total execution time for your entire application, even if the hash is computed much faster, because other concurrent activities may be stalled until the hash computation is completed.
Surprisingly, this is incorrect. Here are some single-threaded measurements on a CPU with SHA-NI: https://bench.cr.yp.to/results-hash.html#amd64-icelake. What you see there is that BLAKE3 can take better advantage of SIMD parallelism than other hashes, and the C and Rust library implementations do this by default. Multithreading isn't enabled by default in the library APIs, but if you do use it (and you have enough input to feed it) the benefits are multiplicative. The b3sum CLI does use multithreading by default.
> only 1 cryptographic hash is faster: BLAKE3
SIMD implementations of KangarooTwelve are also about as fast as BLAKE3, given enough input.
In my tests on AMD Zen 3, BLAKE3 is much faster only when multi-threaded and that should be true on any CPU where the SHA instructions are well implemented.
You are right that it is a little exaggerated to say that BLAKE3 is the only hash with these properties.
Using a similar construction with BLAKE3 to allow parallel computation, instead of using the traditional Merkle–Damgård iterated construction, like SHA-1, SHA-2 and many others, it is possible to design many other fast hash algorithms and there already are many such experimental hashes.
What I have meant was that BLAKE3 is, for now, the only one that has both a freely available and easy to install high-quality implementation, and it is based on a theory that has been studied long enough to have confidence in it.
There are many other hashes that are candidates for being useful fast cryptographic hashes, but it is likely that a few more years are needed to trust them enough.
BLAKE3 is not multithreaded by default.
First, I am not sure the data on most in-use hardware (e.g. EC2 m5/c5/i3en etc ...) supports your conclusions. xxHash is faster than crypto hashes always and BLAKE3 single threaded is faster on every Intel machine I've come across in wide deployment. I hear similar arguments around CRC-32 and to be frank it just isn't true on most computers most people run things on.
Second, many languages don't properly use the hardware instructions and if they do they often don't use them correctly. For example, Java 8 has bog slow SHA-1, AES-GCM and MD5 implementations, and switching to Amazon Coretto Crypto Provider (which is just using proper native crypto) was able to speed SHA/MD5 up by 50% and AES-GCM by ~90% on a reasonably large deployment (although the JDK wasn't using proper hardware instructions for AES-GCM until Java 9 I think it is still slower even after that).
That being said, like I disclaimed at the top of the benchmark your particular hardware and your particular language matters a lot.
[1] https://github.com/corretto/amazon-corretto-crypto-provider/...
With orjson, encoding produces a bytes object instead of a string object, and when you're writing to a file that avoids a bunch of extra memory management overhead and a str.encode() when the text-file wrapper converts that string to a bytes behind the scenes. So that interface change is quite a big performance win over and above just having the faster JSON encoder.
It's not in the standard library, can't just import it and expect it to always be there