Zstandard v1.4.1
github.com
github.com
- By default, the Zstd binary has a lot of additional functionality that gzip doesn't. For example: Zstd includes benchmarking capabilities, the ability to train compression dictionaries, support for legacy (pre-1.0) formats, etc. You can strip all that stuff out with compile flags, and the binary gets a lot smaller.
- Zstandard tries to cover a much wider range of compression speeds than zlib does. To do this, the Zstd compressor actually has eight or so different LZ match-finding implementations under the hood (which get inlined in various ways, resulting in >100 actual versions in the binary), and 5+ different entropy encoding implementations.
For really size-conscious use-cases (like decompressing Zstd content in a mobile app), you can get a minimal Zstd decompressor library down to ~30KB.
On a test workload which slings around many large volumes these changes made it approximately 20x faster[1].
[0] https://github.com/concourse/concourse/releases/tag/v5.4.0
[1] https://github.com/concourse/concourse/issues/3992#issuecomm...
I personally use it all the time in Clickhouse tables (the "Yandex Clickhouse" database). I admit that I'm still using "xz" when I focus on hardcore max compression (max compression within an "acceptable" timeframe) when doing specific tests with their focus on max final compression.
It’s funny though that you should ask that question now. Just yesterday I did submit an erratum on the RFC [0], changing a tiny detail of the spec. While it is literally, a breaking change, (1) we have not yet changed Zstd to produce outputs taking advantage of the change, and (2) it is not actually a concern, because all existing Zstd decoders ignore the existing spec in this regard.
We do actually have a simple “educational” decoder implementation [0], but I wouldn’t recommend its use in production. And then there are two re-implementations in Java [1] and Go [2].
[0] https://github.com/facebook/zstd/tree/dev/doc/educational_de... [1] https://github.com/airlift/aircompressor/tree/master/src/mai... [2] https://github.com/klauspost/compress/tree/master/zstd#zstd
$ du -h data.gz
324M data.gz
$ time -p unzstd -k data.gz
real 1.43
user 1.24
sys 0.19
$ time -p gunzip -k data.gz
real 4.85
user 2.28
sys 0.19
It's even faster at decompressing its own format.