Currently, none. ISA-L might use some, I'm not sure. I have also tried to accelerate some crucial algorithms with SIMD and my trial and errors are still inside the src/benchmarks subfolder. But, in all of those cases, a simple std::string_view::find, std::transform, or lookup tables turned out to be equally fast or faster and, of course, is more portable. Even the CRC32 algorithm uses "simple" lookup tables (slice-by-16) and avoids any SIMD while still achieving ~4.5 GB/s per core. However, there is some Intel Whitepaper [0] showing how to use PCLMULQDQ for CRC32, which might not be much faster but it would reduce cache pressure. AVX-512 even has a VPCLMULQDQ.
[0] https://www.intel.com/content/dam/develop/external/us/en/doc...