Zstandard v1.4.0
github.com
github.com
It allows trading lower CPU for gzip-or-worse compression, and you can mix the settings within a single file. This means you can e.g. use the lowest setting (or no compression at all - it supports that too) to append to a file, while occasionally recompressing recent appends into a single block using the highest setting -- so the cost of compression can be amortized
The only petty annoyance with it is ecosystem support - e.g. GNU tar has no option for it, so it's slightly more painful to work with
My understanding is that Oodle generally performs better (less CPU on both ends for any desired compression ratio) than ZStd in pretty much every context,
http://www.radgametools.com/oodle.htm
... assuming you control both compression and decompression and are willing to spend money for it.
(Besides, regardless of the state of benchmarks today, the momentum is clearly with Zstandard, with Facebook, Intel, and the open source community behind it. The basic lossless compression algorithms haven't changed much since the late 1970s. Making compression fast is mostly just a long slog of engineering hurdles, the kind that big companies are very good at doing.)
Oodle and zstd are both very impressive, it's a shame the former is not free.
This argument by hand-wave doesn’t match up with demonstrated progress over the past few years.
Irrespective of the skill or insight of its individual engineers, I would be surprised if a schizophrenic hack-it-with-duct-tape kind of engineering culture like Facebook could keep up with a focused and motivated expert like cbloom over the medium term (though I suppose the latter could conceivably at some point lose interest in the domain and switch to building something else).
Zstandard is likewise deservedly on track to dominate the lossless compression space.
Some folks disagree:
> [Ruby's] memory usage only reduces when using jemalloc 3; memory usage is still high when using jemalloc 5. Nobody knows why, so that makes the choice of defaulting to jemalloc very dodgy.
via https://www.joyfulbikeshedding.com/blog/2019-03-29-the-statu...
We're working to improve zstd support in the ecosystem over time, but this work moves slowly, and it takes a long time for upstream work to make it to users systems, especially LTS systems.
Is there a way to pass options to zstd when used like this?
With a new enough version of tar you can pass a command with arguments to '-I' (e.g. `tar -I 'zstd -19' ...`).
The alternative is piping the output from tar through zstd yourself.
If you want to pass extra options, you can pipe the output to zstd, which is exactly what tar is doing internally.
Good compression speed ratio (use q 1 for brotli lower than 0.6):
tar -I"brotli -q 2" -cvf file.tar.br inputfiles
Decompression: tar -Ibrotli -xvf file.tar.br
It outperformes gzip and is near xz region without a lot of CPU power and is also blazing fast. Really useful if you have e.g. a big PostgreSQL db dump which you want to transfer to your own machine. Examples for Postgres dumping and restoring: pg_dump adatabase | brotli -q 2 > dump.sql.br
brotli -d < dump.sql.br | psql -U postgres dbusernamehereI still find lbzip2 (which is a bzip2 reimplementation with better algorithms and support for multithreading) quite competitive for highly compressible data. Here's quick and unscientific test that shows that lbzip2 (-9) is still both faster and has better ratio than zstd (-12, --long or not) while also using the least amount of RAM (tmpfs, multithreaded compression using 4-core Xeon E5-2603 v1):
$ time lbzip2 -k linux-5.0.8.tar
real 24.21 user 84.87 sys 4.07 maxrss 105472
$ time ~/zstd-1.4.0/zstd -T0 -k -12 linux-5.0.8.tar -o linux-5.0.8.tar.zst-12
real 30.69 user 105.27 sys 0.61 maxrss 942416
$ time ~/zstd-1.4.0/zstd -T0 -k -12 --long linux-5.0.8.tar -o linux-5.0.8.tar.zst-12-long
real 31.28 user 107.90 sys 0.86 maxrss 1532432
$ time xz -T0 -k -2 linux-5.0.8.tar
real 34.40 user 123.59 sys 0.57 maxrss 410192
$ stat -c '%s %n' linux-5.0.8.tar* | sort -n
126382954 linux-5.0.8.tar.bz2
126394210 linux-5.0.8.tar.zst-12-long
128003669 linux-5.0.8.tar.zst-12
131418488 linux-5.0.8.tar.xz
863426560 linux-5.0.8.tar
The only clear advantage zstd has is decompression speed: $ time xzcat -T0 linux-5.0.8.tar.xz >/dev/null
real 17.25 user 17.06 sys 0.17 maxrss 17312
$ time ~/zstd-1.4.0/zstd -dc -T0 linux-5.0.8.tar.zst-12 >/dev/null
real 2.08 user 1.97 sys 0.08 maxrss 27088
$ time ~/zstd-1.4.0/zstd -dc -T0 linux-5.0.8.tar.zst-12-long >/dev/null
real 2.26 user 2.03 sys 0.17 maxrss 535360
$ time lbzcat linux-5.0.8.tar.bz2 >/dev/null
real 10.34 user 33.74 sys 3.53 maxrss 127088(which is also what makes it a near-perfect codec for HDF5, via blosc-hdf5)
EDIT: no, I was wrong, `zstd -T0` is basically the same as `pzstd`.
EDIT: no surprise, as -T0 also enables multithreading.
-mmt=$(nproc) # use all available cores
-ms=off # disable solid archives (compress each file separately)
-m0=lzma2 # lzma2 has better threading than lzma1
-md=64m # dictionary size
-ma=0 # "fast" mode
-mmf=hc4 # hash chain match finder
-mfb=64 # number of "fast bits"
-mf=off # disable filters
The biggest gains are: 1) using all available cores, 2) setting the match finder (the binary tree match finders are terribly slow; I haven't played much with the newer patricia tree match finders), 3) disabling solid archives (this seems to cause 7zip to distribute the work more evenly between cores, though it still may only use a few cores if there are many small files), 4) using "fast" mode (whatever that is, it gives a noticeable performance boost and doesn't seem to affect compression ratio much).Every few years I try zstd and others, and for the data I work with (primarily a mix of json and fixed-width-field binary data), lots of tools beat 7zip out of the box, but they fall short of 7zip with the above command-line options.
Edit: Ahh, I was confused. Neither require a separate training step. Zstd offers an option to do a training step. Both always use dictionaries with a default size that can optionally be changed.
zstd --long=26 -T0
From there you can tune the compression level, or increase the window size up to 2 GB (--long=31). zstd won't beat the compression of xz, but it can compress much faster if you trade off some space. $ time 7zr a -mmt=$(nproc) -ms=off -m0=lzma2 -md=64m -ma=0 -mmf=hc4 -mfb=64 -mf=off linux-5.0.8.tar{.7z,}
real 60.49 user 158.94 sys 3.06 maxrss 8995040
$ stat -c '%s %n' linux-5.0.8.tar.7z
127700475 linux-5.0.8.tar.7z
$ time 7zr e -so linux-5.0.8.tar.7z >/dev/null
real 14.09 user 13.96 sys 0.12 maxrss 282208
Basically:- it took twice the time to compress data even compared to xz -2 (which also uses lzma2 under the hood),
- it is comparable to zstd/bzip2 ratio-wise,
- it used almost 6 times (!) more RAM than even zstd -12 --long,
- it only used about 2.5 CPU cores out of 4 while compressing (which aligns pretty well with your reasoning for using -ms=off).
----
But hey, source code is not that regular. Since you mentioned JSON and fixed-width-field binary data, I decided to re-run benchmarks on 10M lines of nginx access logs: they're way more regular in their structure (repetitive URLs, timestamps, Mozilla/5.0, stuff like that) that might benefit from larger window sizes.
$ time lbzip2 -k access-log-10m.log
real 90.59 user 313.04 sys 18.46 maxrss 117904
$ time ~/zstd-1.4.0/zstd -T0 -k -12 access-log-10m.log -o access-log-10m.log.zst-12
real 77.34 user 277.21 sys 1.55 maxrss 886416
$ time ~/zstd-1.4.0/zstd -T0 -k -12 --long access-log-10m.log -o access-log-10m.log.zst-12-long
real 69.24 user 242.18 sys 1.85 maxrss 1411872
$ time 7zr a -mmt=$(nproc) -ms=off -m0=lzma2 -md=64m -ma=0 -mmf=hc4 -mfb=64 -mf=off access-log-10m.log{.7z,}
real 109.10 user 356.42 sys 4.69 maxrss 9777664
$ stat -c '%s %n' access-log-10m.log* | sort -n
208537395 access-log-10m.log.bz2
231953002 access-log-10m.log.zst-12-long
237566691 access-log-10m.log.zst-12
249412192 access-log-10m.log.7z
3386733539 access-log-10m.log
Now tweaked 7z did better CPU- and time-wise, but it's still behind zst and bz2 on every metric, especially RAM which it requires so much of (literally gigabytes) it becomes impractical in a number of situations. And we needed a pretty regular input (not just some pretty compressible text like source code or Wikipedia dump) to close that gap. So I can't really recommend your suggestion, unless you have some niche input that benefits from that particular set of options (but then, who has time to learn lzma internals and how every option plays with different kinds of input?).----
It's also worth pointing out that zstd also has plenty of options to fiddle with: https://github.com/facebook/zstd/blob/dev/programs/zstd.1.md...
- maximum compatibility (while tolerating low performance) - gzip
- great performance (while tolerating larger files) - snappy
- very good performance with good (not best) compression ratios - zstd
I don't really want to use Any New Shiny Algo to compress some data that might outlive this piece of software, that's why I use gzip very often, because I know I'll always be able to decompress it. But I've been increasingly adopting zstd and snappy for one single reason - they are becoming widely supported within the ecosystems I work in (data processing).
That, to me, is more important than compression ratios and decompression speeds.
Also, for gzip, you might want to consider using its multithreaded version, called "pigz".
A temporary work around in the meantime, I've used `chattr +C` on the directories I want to be exempt from zstd compression, so that grub can read those files.
zstd -d < compressed.zst | pv > /dev/null == ~330 MB/s
For comparison, pixz with the same data using 32 cores: pixz -d -p 32 < compressed.xz | pv > /dev/null == 1.15 GB/s
Granted, zstd is far, far more efficient per core, but there are plenty of workloads where I can afford to use a lot of cores for decompression. Also pixz still compresses slightly better than zstd -19, but I'd be willing to trade that for more efficient decompression if I could still have the option of really fast decompression using multiple threads.Note also that with this particular data, I'm seeing a compression ratio of only about 4.3:1 using zstd -19. I can imagine that zstd would use less CPU when decompressing if the ratio was higher.
Anyway, for a test file I get ~125% CPU with pzstd -d, so it is able to do more work, and slightly decreases time.
Decompress zstd -d
real 0m19.125s user 0m10.888s sys 0m2.575s
Decompress pzstd -d
real 0m14.819s user 0m14.017s sys 0m4.163s
pzstd is now obsoleted by zstd -T0, but it offers multithreaded decompression for files compressed by pzstd (it will still be single threaded for files compressed by zstd).
Though anecdotal, certainly.
Brotli dominates HTTP compression. Zstd just got its RFC approved a few months ago, but Brotli has been present in browsers for years.
However, zstd is more widely adopted everywhere else, especially in lower level systems. Zstd is present in compressed file systems (BtrFS, SquashFS, and ZFS), Mercurial, databases, caches, tar, libarchive, package managers (rpm and soon pacman). There is a pretty complete list here https://facebook.github.io/zstd/.
Again, I'm biased because I know almost everywhere where zstd is deployed, but not everywhere that Brotli is.
You just cannot compare two things with a commutative function, and that blog post is based on an assumption that you can. Math just doesn't work like that.
Brotli's fastest compression modes are 3-5x faster than zlib's fastest modes. For every gzip quality setting there is a brotli setting that is both faster and more dense than that gzip setting.
That 60 us saving comes with a cost: you'll be transferring 5 % more bytes, which can cost you a hundreds of milliseconds. Brotli is also more streaming, so you get your bytes our earlier during the data transfer. This allows for pages to be rendered with partial content and further fetches to be issued earlier.
Zstd supporters have used comparisons against brotli where they compare a small window brotli against large window zstd. This makes it seem like zstd can compete in density, too, but that is just apples to oranges comparisons.
Zstd has an advantage if you don't have the CPU to compress at the maximum level, since zstd is generally faster than Brotli at the lower levels.
Even still, for web compression, Brotli has the advantage of already being present in the browsers, so you're betting off using Brotli for web compression as it stands today.
Zstd can be significantly slower in medium levels -- you just need not to be tricked to apply compressors at different window sizes. Zstd changes window sizes under the hood. With brotli you need to explicitly decide about your decoding resource use.
If you use the same window size and aim for the same density of compression, brotli actually tends to compress faster in the middle qualities, too.
Some of the obvious competitors are brotli and snappy.
Here some comparison:
https://quixdb.github.io/squash-benchmark/
http://www.mattmahoney.net/dc/text.html (I wonder if anyone has an updated picture with Pareto front for these numbers.)
Brotli with large window can do around 199M for the 1G text corpus: https://groups.google.com/forum/m/#!topic/brotli/aq9f-x_fSY4
Here a large window aggregate result view: https://encode.ru/threads/2947-large-window-brotli-results-a...
There's a 7zip fork that includes Zstd support, but it can only put Zstd inside a .7z container, which doesn't appear to work with any other tools.
In the Windows world, archiving and compression are usually in a single file type (.zip, .rar, .7z). Zstd follows the unix style where it can't directly compress folders of files, they need to be in a tar (or other archive format) first. This isn't really an issue on Linux, since Zstd support is built into tar, which ships on pretty much every system.
You could make uncompressed zip or 7z files and compress that independently as a zst file, but that's a bit baroque compared to just using tar. :)
7-Zip does often seem bent on supporting everything, I imagine some day in the future it'll support zstd at least as an independent archive, if not extending the Zip and 7z formats as well.
- https://github.com/mcmilk/7-Zip-zstd
The changes to the 7-Zip file format were discussed upstream, and upstream agreed to not tread on the magic values:
- https://sourceforge.net/p/sevenzip/discussion/45797/thread/a...
But ultimately the patches were not upstreamed. It's now onto its second developer:
- https://sourceforge.net/p/sevenzip/discussion/45797/thread/6...
Images and video sure, but everything from html, just, css, svg isn't compressed.
In fact compression is critical to modern SPA frameworks to keep initial download times lower.
Also this post has little to do with http compression. ZSTD is used it many other circumstances.