So a test between zstd and brotli would show brotli in a poor light if it used a mixed corpus, but a test between zstd and brotli on a web corpus would give an advantage to brotli...
So a test between zstd and brotli would show brotli in a poor light if it used a mixed corpus, but a test between zstd and brotli on a web corpus would give an advantage to brotli...
I have a feeling that the dictionary was designed with the specific goal of performing well on a specific corpus similar to the Large Text Compression Benchmark[1]. It has quite a few words and phrases that I'd associate with Wikipedia's "house style".
- 9216 phrases total
- 5857 (63.5%) pure ASCII phrases, mostly English with a few Spanish words thrown in
- 1372 (14.8%) code fragments -- mostly HTML, CSS, and Javascript
- 1027 (11.1%) CJK (Chinese, Japanese, and Korean, and mostly the first two) phrases -- it's very hard to tell Chinese and Japanese apart in this context; I didn't try
- 158 (1.7%) phrases containing extended Latin-1 characters (nearly all Spanish words)
- 303 (3.3%) Cyrillic script (probably Russian) phrases
- 322 (3.5%) Arabic phrases
- 172 (1.9%) Devanagari script (Hindi) phrases
Plus a few miscellaneous other scripts and generally unclassifiable content.
Because of its seemingly haphazard dictionary, I wouldn't rule out zstd outperforming brotli, if trained on a good dataset.
Uh ? What format of data was this ?
I did a pretty large test (2gb+) on OBJ/STL 3d data and brotli compressed within ~5% margin of lzma, and this holds true on other binary data I've compared.
It also compressed better than zstd (as in compression ratio) on the same data at their highest respective compression settings:
bro -quality 11 -window 24
zstd --ultra -22
So I find it baffling that you find brotli to be poor for any binary data, could you share the data in question ?
I understand that brotli is incredibly slow as compared to LZMA when using these high ratio settings (q11, w24), so slow as to be impractical in production even in a write-once, read-many scenario if you have any non-trivial amount of data being regularly produced. I do not want to have a farm of machines just to handle the brotli compression load of our data sets just because it is 10x slower than LZMA.
Compression time is indeed the achilles heel of Brotli, which is why it's something I would only use for compress once (and preferably decompress very often) scenarios.
Compared to lzma at it's best compression setting for this particular data (-mx9 -m0=LZMA:d512m:fb273:lc8), brotli took 6 minutes and 4 seconds to compress, while the same data took 1 minute and 47 seconds for lzma.
On the other hand, brotli decompressed the same data in 0.6 seconds, while it took lzma 2.2 seconds.