LZMA compression (used in 7-zip, XZ, LZMA2) is known as one of the best, but it has a noticeable drawback - it's slow. I tried to improve the decompression speed by removing excessive branching at decoding of every bit from the compressed data.
Decompression speedup from this patch largely depends on the compression ratio, more ratio - less speedup. Compressed text, such as source code, gives the least speedup.
That's the result from my Skylake, compiled with GCC. Please help me with testing on different x86 CPUs.
x86 (32-bit) - should work, but haven't tested yet. Compiled with Clang should work as well.