Zstandard – Real-time data compression algorithm
facebook.github.io
facebook.github.io
some cool references:
https://www.youtube.com/watch?v=uXtmN9fE01k
https://th.if.uj.edu.pl/~dudaj/
https://demonstrations.wolfram.com/DataCompressionUsingAsymm...
https://encode.su/threads/2078-List-of-Asymmetric-Numeral-Sy...
well, impressive
The thinking seems to be due to how CPU intensive it is, Brotil is not favoured for on-the-fly compression[2].
If you're really desperate for it though, you could try [this extension][3].
[1] https://caddyserver.com/docs/caddyfile/directives/file_serve...
Other webservers support it, mainly Nginx. But since Caddy is written in go, they need a native go port which doesn't exist yet.
Yeah, hence my "really desperate" qualifier. I probably should have been more explicit.
https://quixdb.github.io/squash-benchmark/unstable/?dataset=...
Perhaps, this common misunderstanding is due to the fact the slowest level (-11) is the default one in the brotli command line utility?
It's also faster than gzip, and gzip was used for "just in time" compression on the web for ages, so I don't see how brotli can be a problem for this use case. (I assume by "encrypt" you mean "compress".)
https://github.com/facebook/zstd#the-case-for-small-data-com...
zStd also supports pre-trained dictionaries, it is particularly good for specific small JSON, for example. It won't be that helpful for relatively large HTML files.
Common reason being someone forgot to check how many bytes read() actually returned and just assume it filled the buffer.
Robust data formats have double protection against this class of bugs with unambiguous SOF and EOF boundaries that the consumer can assert. I guess a sanity-byte once every X Megabytes wouldn’t hurt either.
I still don't understand the use case limitation, or exactly where this is a potential issue. Why is the data being interrupted? Is this some kind of raw UDP stream use case?
zstd 1.4.5 -1 2.884 500 MB/s 1660 MB/s
zstd 1.4.5 --fast=1 2.434 570 MB/s 2200 MB/s
zstd 1.4.5 --fast=5 2.178 700 MB/s 2420 MB/s
lz4 1.9.2 2.101 740 MB/s 4530 MB/s
3% larger size for a nearly 2x decompression speed is a no brainer for lots of uses.Anecdotally I have also gotten much better decomp speed on snappy on previous generations of CPUs staying within ~10% compression ratio, though I had not benchmarked/tuned extensively nor reconfirmed in the past ~2 years. (Probably this can also vary a lot by implementation, which you're often stuck with based on some other library choice.)
For my use case, any extra benefits of zstd are far outweighed by the extra cost of downloading zstd itself in the browser.
LZ4 is smaller, faster and much more simple.
zstd also seems to be optimizing much more, recently, for the large stream of data use cases, versus the fixed-size or smaller records you might get in, say, databases, or filesystems, so I think it's actually becoming less generally useful over time of late.
And if you're going to play with Snappy, you might find S2, which was linked on HN relatively recently, interesting. [1]
I imagine that xz might be good enough at making its own dictionary that what it ends up with wouldn't be much different from what I'd make manually, so I could see a decent chance that the improvements would be minimal at best. Has anyone done any experimentation along these lines?
Depending on your needs you could also look at the MySQL clone plugin, which gives you a simple way to clone the data dictionary as physical backup, avoiding the textual representation, but has its restriction. https://dev.mysql.com/doc/refman/8.0/en/clone-plugin.html
Disclaimer: I work at the MySQL development team at Oracle
zstd 1.4.5 --fast=5 2.178 700 MB/s 2420 MB/s
lz4 1.9.2 2.101 740 MB/s 4530 MB/s
Read the LZ4 format, it's very simple and educational, I'd say beautiful. It fits in half a screen of text.
(see, e.g., https://news.ycombinator.com/item?id=32529412#32531734)
For some applications it matters how fast one operation completes, but for others, it's much more relevant how much CPU time it consumes, so if zstd needs 1 second on 8 cores for something gzip does in 8 seconds on 1 core, it would be no benefit at all.
I can see why you would want it to be efficient, economical in bandwidth, and fast enough in all scenarios. But why does it need to be faster in situations where it is already quite fast?