In my scenario, I'm decompressing from HDD and piping the decompressed data directly to another process. I may be limited by disk sometimes, but often the data is already in the filesystem cache. To give a concrete example:
zstd -d < compressed.zst | pv > /dev/null == ~330 MB/s
For comparison, pixz with the same data using 32 cores:
pixz -d -p 32 < compressed.xz | pv > /dev/null == 1.15 GB/s
Granted, zstd is far, far more efficient per core, but there are plenty of workloads where I can afford to use a lot of cores for decompression. Also pixz still compresses slightly better than zstd -19, but I'd be willing to trade that for more efficient decompression if I could still have the option of really fast decompression using multiple threads.
Note also that with this particular data, I'm seeing a compression ratio of only about 4.3:1 using zstd -19. I can imagine that zstd would use less CPU when decompressing if the ratio was higher.