RAD's ground breaking lossless compression product benchmarked
richg42.blogspot.com
richg42.blogspot.com
I can't put my finger on it, but something about the blog post bothers me. The "groundbreaking" compression suite is closed source and commercial, and even the data set used for testing in the post isn't public. Con Kolivas (ck), the author of lrzip, commented and inquired about the corpus but the blog post author says he "can't release it". I dunno, it just seems like a weird exchange to me.
[0] http://www.radgametools.com/oodlewhatsnew.htm [1] https://news.ycombinator.com/item?id=11583898
Of course he can't say in public that he's kept copies of any of those assets. They belong to his previous employers.
Yep, same here. Way too much excitement for something that shows no "groundbreaking" advantage over other codecs based on his own benchmark results.
Eh..Why should anyone in the world care that a compressor can compress test data well? Except vaguely, it might serve as a motivator to investigate the use of said compressor for their own purpose. But then again, this is simply marketing, for developers.
To me, its like saying "the new compiler version is faster" on this collection of source files. Personally, I couldn't care less what the test set is, if the compiler is not faster for my own use case.
So yeah, if anyone is interested, they could probably get an evaluation copy, try it out on their own data and decide for themselves.
Pretty crude heuristic, but seems to work with very few false positives!
Sadly, a compression benchmark that doesn't detail the compiler, hardware, corpus, or even version of the libraries used is not very compelling. If others cannot reproduce your results, the results aren't of much value beyond marketing.
Stuff like the compilers used for Oodle are detailed there, and a bunch more benchmarks are published using (I believe) open corpora.
And these high end benchmarks for bigger data: http://mattmahoney.net/dc/silesia.html http://mattmahoney.net/dc/text.html http://mattmahoney.net/dc/10gb.html
Seems like LZ4 is still unbeatable if you need as fast compression as possible. Think kernel memory page compression (to avoid swapping pages to disk) and compressing data destined to a high bandwidth sink. I'm just guessing, but probably LZ4 has also less power consumption.
If your output device can do over ~100 megabytes per second, LZ4 seems to win easily.
Optimal compression algorithm selection seems to depend on capacity and bandwidth constraints of output device.
$ memo expensive-with-lots-of-output | memo -s expensive | less
I noticed that lz4 (the default (de)compressor memo uses) was the bottleneck. ("memo -s" absorbs standard input and calculates a checksum over it and the command arguments, if that checksum is in the cache, it knows it can just output the cached value instead of running the command).So I started looking around if there was nothing better around. From the Squish benchmarks posted elsewhere on this article [2], I noticed that zstd(1) seemed to make the right trade-offs. I required compression speed > 50MB/s, decompression speed at least as fast, and as high a compression ration as possible.
zstd(1) seems to work brilliantly and is very fast, it was easy to compile too, just running make in its directory was sufficient for me. I love those kinds of projects. I've made it the default over lz4. What gave me even more confidence in this change is that zstd is by the author of lz4.
[1]: https://github.com/aktau/dotfiles/blob/master/bin/memo (currently a version that doesn't prefer zstd(1) yet, I need to git push).
It probably depend on how much bandwidth you got.
[1] http://cbloomrants.blogspot.com/2016/07/introducing-oodle-me...
[2] http://cbloomrants.blogspot.nl/2016/07/oodle-selkie.html
However, the compression is nothing compared to Pied Piper's algorithm.