Zstandard v1.4.7
github.com
github.com
If anyone from Facebook is reading, please look at making zstd seekable. There is already experimental support but it needs to be standardized. This could open up really interesting new use-cases [see my comment linked above], where currently we have to use much slower xz/lzma instead.
I'm a sysadmin so this has worked it's way into every corner of my servers, from backups, to log compression, and more. Most of my servers run FreeBSD, and with the merge of linux/FreeBSD ZFS I am looking forward to the integration of ZSTD into ZFS so I can use that for on-the-fly Zpool compression.
I can safely say that ZSTD has saved me TB of space over LZ4 & GZIP.
I've been a zstd convert for a few years now and haven't settled on any particular ideal setting. Sometimes the very high compression levels (-19) produce a worse result, sometimes they kill it. I think most of the time I'm using -9.
I've seen cases where '--ultra --long -22' was bringing size down by 10% or more compared to the normal settings, and also cases where changing the defaults at all produced significantly lower compression.
(On the dataset used for the figure, zstd -9 was about the same MB/s as gzip -4, and compressed at ~3.5x to gzip's ~3.0x.)
https://en.wikipedia.org/wiki/Block_cipher_mode_of_operation
To put some numbers to this, on the machine (>10yo laptop) I'm typing this on:
trivial random .wav file
LZMA ~20 MB/s ~2 MB/s ~1 MB/s
zstd ~500 MB/s ~250 MB/s ~70 MB/s
zlib ~40 MB/s ~13 MB/s ~6 MB/sI would love if xz had a zstd filter and a filter that adds/processes some error correction code. But with what's there right now I think it's fair to treat it as a pure lzma2 container format.
Side by side comparison right now between a full system file tree backup of a mediawiki server, the .tar.xz file is about 1.8GB, the .tar.zstd file is 2.15GB. The same thing in traditional .tar.gz format is obviously even larger.
From the research I've done on common uses of zstd it's much better suited to applications where very high speed of compression and decompression is a major requirement, but absolute smallest file size is not. Typically as a replacement for gzip.
It does have a performance advantage in that it by default uses both CPU cores, while a default tar with -J option will use xzip on only one CPU core.
> zstd and xz trade blows in their compression ratio. Recompressing all packages to zstd with our options yields a total ~0.8% increase in package size on all of our packages combined, but the decompression time for all packages saw a ~1300% speedup.
Often times persons on very limited bandwidth connections (even a Hughesnet or Viasat consumer grade satellite link in a very rural area of the USA) will have more than sufficient CPU to decompress xz in a reasonable amount of time. Such as a five year old core i5 desktop computer.
In environments where it can be expected that the bulk of users are on very fast broadband connections, zstd and slightly larger file sizes have an advantage.
Take for example ~13.5 GB. This is a 6 hour download at your 5Mbps figure. An extra 0.8% is less than 3 minutes on that 6 hour download.
If it might take ten hours to transfer 900MB in some unusual environments, that's a big savings.
Also there is a degree of asymmetry, in that it's not so difficult for the system creating the xzip archive to have a lot of CPU (four CPU cores out of 12 or 16 on a dual socket current generation xeon system, or a threadripper). Since xzip is much more time consuming to create than to extract.
Or for any other environment where WAN connectivity is extremely limited in throughput.
If I had to guess, since this originated from within Facebook, they're using the feature for user definable dictionary to maintain a large corpus of sample, representative user uploaded content.
Which would accomplish the purpose of tuning its performance for their own specific internal usage.
The ability to define your own dictionary is useful beyond simply large indexable datasets. Anything application-specific can usually benefit from user-defined dictionaries.
[1] http://fastcompression.blogspot.com/2013/12/finite-state-ent...
For someone that's a web developer that's never done game development at all, do you have any recommended resources for learning?
There's a ton of resources out there for learning engines - Unity, Unreal, Godot, etc. I went the low-level route and built a custom engine, which was fun and taught me a lot about programming. No single resource comes to mind - requires piecing together a lot of different things.
The canonical paper on ANS is indecipherable to me (i honestly think it's poorly written, it's not just me, although it is partly me!), but this seems like a clearer explanation (although i still don't entirely get it):
https://bjlkeng.github.io/posts/lossless-compression-with-as...
zstd --patch-from=old new -c | zstd -d --patch-from=old
Using the library this can be achieved by using the old file as a dictionary. If you have very large files (MB-GB range) you will also want to enable long range mode, otherwise zstd might not be able to find the repetition. --patch-from on the CLI will manage this automatically.
The release notes from zstd-1.4.5 [0] contain some details about the speed/size tradeoff of "patch-from" mode, compared to other delta engines. There is also a wiki page detailing how to use zstd as a delta engine [1].
[0] https://github.com/facebook/zstd/releases/tag/v1.4.5
[1] https://github.com/facebook/zstd/wiki/Zstandard-as-a-patchin...
But that is throughput. How latency (on small files/chunks) is affected I have no idea.
Lacking a clear source for a public dictionary suited to the web corpus, there's not much point in putting it in browsers.
Facebook is using it for its HTTP traffic to its apps though. We've seen significant benefits. The compression ratio is pretty comparable to Brotli, but it decompresses very much faster, which is nice when the client is a mobile device--lower latency, longer battery life, etc.
There's discussion on this topic here: https://github.com/facebook/zstd/issues/1355