[1] http://fastcompression.blogspot.com/2016/06/zstandard-reache...
[1] http://fastcompression.blogspot.com/2016/06/zstandard-reache...
Brotli compresses usually more, and quite a lot more on shorter files. Try with cp.html, sparc sum or xargs.1. Brotli compresses these 9-19 % more densely on the following benchmark:
https://quixdb.github.io/squash-benchmark/unstable/
Note, that on this benchmark brotli is always limited to the 4 MB sliding window. Other algorithms are run with wider windows, too. This will make brotli seem worse on large files (4+ MB).
And especially thanks for the clarification on performance relative to input file size and sliding window size.
Come on, this is not serious.
Brotli's fastest compression algorithm is still significantly slower than zstd. And more importantly, it compresses _much worse_.
For a 3rd party evaluation, one can try [TurboBench](https://github.com/powturbo/TurboBench) or even [lzbench](https://github.com/inikep/lzbench) which are open-sourced. Squash introduces a wrapper layer with distortions which makes it less reliable, and more complex to use and install, quite a pity given the graphical presentation is very good. I'm interested in speed, and in this area, all benchmarks point in the same direction : for a given speed budget, Zstandard offers better ratio (and decompresses much faster).
In any case, neither the compression ratios nor the speed of large-file compression necessarily say much about small file performance. There's just much more context to search in a 100MB file than there is in a 10KB file.
Having said that, there's no reason to assume brotli is better for small files; there's just no way to tell given the links you provide.
I'm not affiliated with nor use neither zstd nor brotli.
I get a test file by: wget https://web.archive.org/web/20151222062543/http://www.micros...
The test file is 267253 bytes.
$ ./lzbench -ebrotli,0,1,2,5,7,9,11/zstd,1,22 testfile
brotli 0.4.0 -0 compresses 783 MB/s and decompresses 809 MB/s
zstd 0.7.1 -1 compresses 586 MB/s and decompresses 1691 MB/s
brotli 0.4.0 -7 compresses 57 MB/s, decompresses 873 MB/s to 28185 bytes
brotli 0.4.0 -11 compresses to 25413 bytes
zstd 0.7.1 -22 compresses in 4.01 MB/s to 28363 bytes
Of course it is an unfair example because of the static dictionary that brotli uses, but it is not a pathological example: Thai is not part of the static dictionary. The numbers are on a i7-4790K@4.00 GHz.
Brotli's fastest compression is faster than that of zstd, at least as shown with lzbench and this file. Also brotli wins in compression density. In this file the win is 10.5 % less bytes for brotli -11 than for zstd -22.
It is certainly a favorable ground for Brotli. Brotli claims an advantage in html compression, thanks to its integrated specialized dictionary. The real pb though is the suggested conclusion that these favorable results are broadly applicable everywhere else. That's a terrible suggestion. We need more examples, not just "html files" which happen to be Brotli's best case.
> brotli 0.4.0 -11 compresses to 25413 bytes > zstd 0.7.1 -22 compresses in 4.01 MB/s to 28363 bytes
Why you don't disclose the compression time of brotli ? Of course it does matter : everyone understand that an algorithm that spend 10x more cpu has the budget to compress more.
> brotli 0.4.0 -0 compresses 783 MB/s and decompresses 809 MB/s > zstd 0.7.1 -1 compresses 586 MB/s and decompresses 1691 MB/s
Here, you don't disclose the compression ratio of both algorithm, implying they are equal. By such standard, LZ4 is probably the best : it's so much faster ! Of course, they do not compress the same...
I was initially thrilled at your detailed answer, but now, quite frankly, I feel cheated. Grossly so.
This is really disappointing. I was so much vexed that I decided to run the tests myself.
Downloading and using __the same html file__, the same lzbench, same library versions, just a different computer and compiler, here is what it produced :
| Algo | compressed size | compression speed | decompression speed | | --------- | --------------- | ----------------- | ------------------- | | brotli -2 | 36223 | 220 MB/s | 670 MB/s | | zstd -1 | 36655 | 480 MB/ | 1400 MB/s | | brotli -1 | 38292 | 360 MB/s | 650 MB/s | | brotli -0 | 41141 | 560 MB/s | 610 MB/s |
__Conclusion__ : brotli -0 is indeed fast, faster than in my previous tests. It seems to be tuned to reach this objective, but throw away a lot of compression ratio to get there.
Consequently, brotli -0 is not comparable to zstd -1, it takes brotli -2 to produce an equivalent compressed size . By that time though, zstd is much, much faster.
Which is exactly the question I'm trying to get answers to : which algorithm compresses better for a given speed budget ? That's what matters, at least in our datacenter.
I'm not interested in ultra slow mode, but while at it, I wanted to complete the picture with the missing compression speed of brotli - 11. It produced : zstd -22 : 2.95 MB/s brotli -11 : 0.53 MB/s
So that's > 5x difference. It surely helps to reach better compression ratios.
I also wanted an answer to "by how much the dictionary helps ?".
Fortunately, TurboBench can help, thanks to a special mode which turns off dictionary compression. By using it on the very same sample, brotli -11 compressed size increases from 25413 to 26639 bytes. 5% larger, clearly not negligible. Still good, but it cuts the advertised size difference in half.
Anyway, clearly I feel disappointed to have to redo the tests myself, because some inconvenient results were intentionally undisclosed (or not produced). This really undermines my trust in future publications.
That learned me something : trust only benchmark done by yourself. And now, I should probably benchmark even more ...
| Algo | compressed size | compression speed | decompression speed |
| --------- | --------------- | ----------------- | ------------------- |
| brotli -2 | 36223 | 220 MB/s | 670 MB/s |
| zstd -1 | 36655 | 480 MB/ | 1400 MB/s |
| brotli -1 | 38292 | 360 MB/s | 650 MB/s |
| brotli -0 | 41141 | 560 MB/s | 610 MB/s |If you are interested at zstd 0.7.1 -22, you can reach the same compression density with brotli 0.4.0 at quality setting -7 (at least for this file with the static dictionary). Then you are comparing a brotli's compression speed of 57 MB/s to zstd's 4.01 MB/s. Brotli is 14x faster at this compression density.
For decompress-once roundtrip at this density, brotli achieves 53 MB/s, and zstd 4 MB/s. Brotli's roundtrip is 13x faster.
In decompression brotli is 2x slower on Intel, but decompression times at 800+ MB/s are going to be negligible in most use (think < 1 % of cycles in your datacenter), if the data is parsed/processed somehow afterwards.
Brotli's entropy encoding is simpler (no 64-bit operations), and because of this on 32 bit arm the decompression speeds of zstd and brotli are about the same.
I acknowledge that there can be use cases where zstd 0.7.1 can be favorable to brotli 0.4.0 -- particularly those where a 32+ MB file is compressed at once and the 150-500 MB/s compression speed range, but even this simple compression test shows that brotli can compress significantly (5-10+ %) more with the higher quality settings.
I left it out because zstd didn't produce comparable output size.
I showed the compression and decompression speed at brotli at quality 7 that already compressed more densely than zstd at maximum setting. Of course it is a somewhat technically flawed comparison, because brotli gets an advantage for this kind of data from its dictionary, but I invite you to test using another set of files. (The three small compression benchmark files discussed earlier in this thread show similar trend.)
Right now, we're using LZSS because the decompressor code is tiny (hundreds of bytes), and scratch memory is in the single-digit kilobytes.
Any suggestions? All benchmarks I can find online are focused on desktop usage (where comp/decomp memory use is either not mentioned or in the megabyte range) or web streaming usage (where total comp+decomp time is important. I don't care (within reason) about comp time)
I even have some public code for it, there's Python module [2] that will adequately address your wish for mediocre offline-friendly performance (heh) and a streaming C decompression library [3].
The runtime memory usage for the decompressor are tiny (which is good, since from my perspective your system sounds huge!) so that should be fine at least.
Very interested in any feedback you might have, of course.
[1] https://en.wikipedia.org/wiki/LZJB [2] https://github.com/unwind/python-lzjb [3] https://github.com/unwind/lzjb-stream
http://www.cbloom.com/rants.html (this is his blog, you’ll have to skim back through the last year of posts). But e.g. check out this very recent one http://cbloomrants.blogspot.com/2016/05/ps4-battle-miniz-vs-...
http://www.radgametools.com/oodle.htm http://www.radgametools.com/oodlewhatsnew.htm