Stop using gzip
imoverclocked.blogspot.com
imoverclocked.blogspot.com
On the other hand, having to deal with support requests from users who don't have any decompressor other than gzip will cost me both users and my time. Some complicated "download this one if you have xz" or "here's how to install xz-utils on Debian, on RHEL, on ..." will definitely cost me users, compared to "if you're on a UNIXish system, run this command".
From a pure programming point of view, sure, xz is better. But there's nothing convincing me to make the engineering decision to adopt it. The practical benefits are unnoticeable, and the practical downsides are concrete.
Well, at the time it was released, people were making much the same arguments (with kittens). It compressed so much better than gzip, no reason to use the obsolete gzip format and tools, etc. And some of us jumped on the hype bandwagon and started recompressing our data, only to find out afterwards that bzip2 is now the new thing and the format is not only obsolete, but also patent-encumbered and in general needs to be phased out.
From a long-term perspective I'm fine with gzip. At least I know that I'll be able to open my data in 10 years time, which is not the case with bzip-0.21. The jury is still out on "xz", in my opinion.
If you're a developer, gzip is simply the best option. It's not the best, but it's good enough and it's safe.
* The faster compression made a difference of about 100 milliseconds to a user experience lasting minutes.
* The better compression made a different of about 1 second to most users.
* The change of compression algorithm would take time away from engineering teams, and ultimately introduce bugs.
So in the end it was sidelined until other fundamental changes (file format etc) made it able to be coat-tailed into production.
Same story as yours: don't fix what ain't broke. Tallest working radio antenna in the world -> longest lying broken antenna in the world.
Mostly Mac, but iirc some Linux users were confused as well. We switched back to gzip because everyone knows how to use it.
tar -xf some_file.tar.xz
tar -xf some_file.tar.gz tar -xzf file.tar.gz
To be honest, I only started doing -xf a year ago. I was used to -xjf and -xzf. tar -xzvf file.tgz > /dev/nullAlso works with non-tar formats like Zip, RAR and 7z: https://www.freebsd.org/cgi/man.cgi?query=libarchive-formats...
$ tar -xf foo.tar.xz
tar: End of archive volume 1 reached
tar: input compressed with xz
$ tar -xf foo.tar.gz
tar: End of archive volume 1 reached
tar: input compressed with gzip; use the -z option to decompress itNot letting untrusted input automatically increase the attack surface it's exposed to is a feature.
How is that a feature? The user's explicitly asking for this.
This feature reminds me of vim, that suggests closing with ":quit" when you press C-x C-c (i.e. the keychord to close emacs). It knows full well what you want to do and even has special code to handle it, but then insists to hand you more work.
Upon receiving a C-c, it does not know full well what the user wants to do.
When vim receives a C-c from you (or someone who just stumbled into vim and doesn't know how to exit) the user wants to exit.
When vim receives a C-c from me, it's because I meant to kill the process I spawned from vim, and it ended before the key was pressed. I very much do not want it to quit on me at that point.
Showing a message seems the best compromise.
I don't really care what vim does, that's a different argument. There have been many vulnerabilities in gzip, and in tar implementations that let untrusted input choose how it gets parsed, those vulnerabilities might as well be in tar itself.
The workaround that the parent is talking about is usually "get LAME from a different distributor", which is still done by Audacity and others.
Truth.
Chamber music in an echo-y cathedral. With bad encoding, you can hear a noticeable difference in the length of time the reverberations are audible and and the timbre of those reverberations cab be quite different. With lots of acoustic music, the "accidental beauty" produced by such effects can be quite important.
Finding this convinced me to re-encode my music collection in 320kbps MP3 for anything high quality, and algorithmically chosen variable bitrates for lower quality recordings -- usually around 160 kbps. That was quite a number of years ago, though. I'd probably use another format today.
On the other hand 128 kbps AAC is transparent for almost any input. AAC is supported abou everywhere where mp3 is. The quality alone should be convincing. The smaller size make the usage of mp3 IMHO insane.
OTOH "the scene" still does MPEG-2 releases I think.
> Really at 320kps you are entering the realm of fantasy if you think you can hear any difference.
It depends on the encoder, the track, your equipment, and how good you are at picking out artifacts. Some people do surprisingly well in double-blind tests, though I doubt anyone can do it all the time on every sample.
That people have been doing this for many years is one of the big reasons modern encoders are so good - they've needed tonnes of careful tuning to get to this point.
Hydrogenaudio listening tests [1] are studies by volunteers, but they focus on non-transparent compression. Anyway, it also illustrate aswell how bad mp3 is.
[1] http://wiki.hydrogenaud.io/index.php?title=Hydrogenaudio_Lis...
And it was reproducible on all my computers. I couldn't use my opus collection at all because either the encoder was broken or the playback was broken.
I need to do it again at some point, when I have 8 hours to transcode everything and try again. See if they've fixed it.
Don't ever do that. If you didn't rip CDs to lossless or buy lossless, you are stuck with the format you've got.
Debian/Ubuntu dpkg supports compressing with xz too -- and it's a hard dependency of dpkg at least as far back as the precise (12.04) LTS release. So I'd say the majority of Linux users already have access to xz.
What about tooling? OSX: tar -xf some.tar.xz (WORKS!) Linux: tar -xf some.tar.xz (WORKS!) Windows: ? (No idea, I haven't touched the platform in a while... should WORK!)
Yep, in the same way it requires libz and libbz2.
Both tend to be part of an essential core of packages required to install a system.
Considering these features:
* Compression ratio
* Compression speed
* Decompression speed
* Ubiquity
And considering these methods: * lzop
* gzip
* bzip2
* xz
You get spectrums like this: * Ratio: (worse) lzop gzip bzip2 xz (better)
* C.Speed: (worse) bzip2 xz gzip lzop (better)
* D.Speed: (worse) bzip2 xz gzip lzop (better)
* Ubiquity: (worse) lzop xz bzip2 gzip (better)
So, xz, lzop, and gzip are all the "best" at something. Bzip2 isn't the best at anything anymore.But despite zpaq being public domain, few people have heard of it and the debian package is ancient, and so the ubiquity argument really does count for something after all.
https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=777123 - get in touch with the new owner of the package if you're interested. It's probably on their Never Ending Open Source To Do List.
Are you implying that xz out-compresses zpaq? Can you supply a benchmark?
Here's one from me - http://mattmahoney.net/dc/text.html - showing a very significant compression ratio advantage to zpaq.
* xz is faster than bzip2 and provides better compression ratio [than bzip2]
* zpaq is slower [than bzip2 and provides better compression ratio than bzip2]
But looks like I'm mistaken? It seems like it can be faster and give better compression ratio than bzip2, can it?
Two points:
(1) It's very, very easy for the best solution to a problem not to simultaneously be the best along any single dimension. If you see a spectrum where each dimension has a unique #1 and all the #2s are the same thing, that #2 solution is pretty likely to be the best of all the solutions. Your hypothetical example does actually make a compelling argument that bzip2 is useless, but that's not because it doesn't come in #1 anywhere; it's because it comes in behind xz everywhere. (Except ubiquity, but that's likely to change pretty quickly in the face of total obsolescence.)
(2) lzop, in your example, is technically "the best at something". But that something is compression and decompression speed, and if your only goal is to optimize those you can do much better by not using lzop (0 milliseconds to compress and decompress!). So that's actually a terrible hypothetical result for lzop.
Heck, zero compression easily wins three of your four categories.
Last week, there was a drive mount that was filling up, rate was roughly 30Gb/hr. The contents of that mount was used by the web application. Deletion was not an option. Something that compressed quickly was needed. And on the retrieval end, when the web app needs to do decompression, seconds matter.
I should probably mention the compression ratio was slightly worse than bz2 (maybe 15% larger archive) but for the 10x increase in throughput I didn't really mind that much. I could actually analyze my data from my laptop!
For compression lz4 took ~22 seconds (~210 MB/s) and I got ~30% compression, gzip -1 took ~56 seconds (~80 MB/s) and I got ~22% compression.
For decompression lz4 gave me 500MB/s while gunzip gave me 300MB/s.
Commands used:
lz4 -cd RS_full_corpus.lz4 | pv | head -5000000 | gzip -1 > test.gz
gunzip -c test.gz | pv > /dev/null
lz4 -cd RS_full_corpus.lz4 | pv | head -5000000 | lz4
lz4 -cd stdin.lz4 | pv >/dev/nullThe biggest problem is using a parser that can do 600MB/s streaming parsing. If you use a command line parser don't try jq even with gnu parallel.
That said, parallel XZ is even better: https://github.com/vasi/pixz
-T, --threads=NUM use at most NUM threads; the default is 1; set to 0
to use the number of processor cores> Multithreaded compression and decompression are not implemented > yet, so this option has no effect for now.
Changing from one compression format to another seems harmless, but it always pays to think carefully about the implications.
[1]: https://www.freebsd.org/cgi/man.cgi?query=xz&sektion=1&manpa...
`pigz` has a similar flag that works reliably, though.
Another way compression formats can win you much more than a 2x space reduction is by supporting random access within their contained files. Gzip sort of supports this if you work hard at it. Xz and bzip2 appears similar (though the details are different). I achieved a 50x speedup with this in real applications, and discussed it a bit here: http://stackoverflow.com/questions/429987/compression-format...
http://lh3.github.io/2014/07/05/random-access-to-zlib-compre...
http://www.ebaytechblog.com/2015/10/09/gzinga-seekable-and-s...
And you are right for embedded! .xz just doesn't work there.
I've also found that on the faster systems, for different uses of mine, when I want the compression to last as little as possible and the total round trip time matters (compression and decompression), gzip -1 gives the best resulting size for the reasonably short time I want to spend.
If it's being downloaded once
If it takes you 60 seconds to download as gz, and 50 as xz. The decompression needs to take less than 10 seconds more for it to be comparable and you've got to be sure that your end users have more memory and sufficient processing power to throw at the task.
Admittedly, in most cases, that isn't much excuse though.
http://backdrift.org/oom-killer-how-to-create-oom-exclusions...
The challenge of sharing internet-wide scan data has unearthed a few issues with creating and processing large datasets.
The IC12 project[1] used zpaq, which ended up compressing to almost half the size of gzip. The downside is that it took nearly two weeks and 16 cores to convert the zpaq data to a format other tools could use.
The Critical.IO project[2] used pbzip2, which worked amazingly well, except when processing the data with Java-based tool chains (Hadoop, etc). The Java BZ2 libraries had trouble with the parallel version of bzip2.
We chose gzip with Project Sonar[3], and although the compression isn't great, it was widely compatible with the tools people used to crunch the data, and we get parallel compression/decompression via pigz.
In the latest example, the Censys.io[4] project switched to LZ4 and threw data processing compatibility to the wind (in favor of bandwith and a hosted search engine).
-HD
1. http://internetcensus2012.bitbucket.org/images.html 2. https://scans.io/study/sonar.cio 3. https://sonar.labs.rapid7.com/ 4. https://censys.io/
Fish shell users can take advantage of the Extract and Compress plugins I wrote, which utilize Pixz if installed: https://github.com/justinmayer/tackle/tree/master/plugins/ex...
7zip is the program you want to handle most everything, with both gui and command line options: http://www.7-zip.org/
Given how radically MS is trying to reform itself to be an open-source friendly company and how ineffectually inoffensive they've been the last 5 years, can we at least try and throw them a bone or two?
I've preferred Mac systems for longer than most of the HN crowd has been alive, so I understand what it's like to feel ignored and in the minority. For years, Mac users were treated as pariahs. The tables have turned, and as someone who has been in your situation, I should have great empathy for your predicament.
And I do, but your tone in general -- and your last sentence in particular -- makes it very hard to empathize. Microsoft used its near-monopoly status to stifle innovation for years, and many of us have figurative scars that will never heal. You seem to think that they are making great strides (while I see them as half-hearted overtures), but either way I'm not about to "throw them a bone." Their decades of misdeeds, in my eyes, will not be expiated so easily.
Perhaps Microsoft will someday be worthy of forgiveness, either from the perspective of morality (e.g., Mozilla) or product excellence (e.g., Apple). Until that day, Microsoft will continue to reap what they have sown, given no more attention than they have earned.
I feel like Apple's has forgotten what was important, created then ruined a market, and lost everything that made it interesting (long before Steve Jobs passed away, by the way). Which cuts all the more deeply because back in the early 2ks they were walking the walk and taking a lot from NeXT's culture of developer friendliness. I grew up deeply invested in Macs and NeXT, which makes the realization painful, but... Apple wants to annihilate maker culture as it monetizes its platform. It's also stopped caring about design on a grand scale, instead appealing to very shallow notions of "visual simplicity".
That's all gone now, and they're consequently useless to me. I'd rather patronize a company currently doing the right thing after a troubled past than pretend a previously aligned company was still there.
It should be very telling that Apple AND Google's flagship hardware announcement of 2015 was something that Microsoft has been doing for years.
And if Microsoft suddenly goes evil again? Fuck them, I'll drop them and move somewhere else. Not Linux, unless the distros pull their act together, but I'm sure a competitor will emerge. Or I'll make one.
It's sort of a rough time for devs right now even as we enjoy unprecedented prosperity and recognition. Big businesses are attempting to monetize and control every aspect of developers.
> Windows: ? (No idea, I haven't touched the platform in a while... should WORK!)
I do think the author missed the ball in not doing basic research for the Windows platform. His point is that people should switch from one compression tool to another. If people on Windows were unable to compress or decompress such files, that would be a huge problem for his argument.
That said, I don't agree with the GP's tangent.
Powershell can run on Linux, too. I've even met a few people who quietly prefer it.
So what, exactly, were you referring to? Shell choices are like editor choices: arbitrary and largely equivalent and without any real meaning or impact on a developer's productivity.
Right and putting that in google, "tar equivalent for windows", immediately nets 5 useful results. You can use tar, or a windows command line variant of 7z or tar, or a gui.
> With out-of-the-box Windows machine, you cannot ssh, you cannot untar, etc.
On an out-of-the-box Linux machine, you generally can't do a lot of things either. It seems particularly ironic that in a discussion about how we shouldn't be using old UNIX tools just because they're entrenched, you then call for compatibility.
> Windows-way of doing things is totally different than *nix culture (OS X, Linux, etc).
Stupid legacy path limits not included, Powershell is in my experience just a superior way to do things. I should maybe restart my blog to talk about that.
But even if we ignore Powershell and windows, your statement is divisive within the Linux community. MANY people prefer shells on Linux that don't adhere to the bash legacy. TCSH and CSH are very popular, to this day. Are they 'not interoperable'?
Everyone's got a big chip on their shoulder about how development tooling "should be." One of the things I've come to realize is how arbitrary, unnecessary, and useless these mores are. They just hold us back.
Minimal distros aside, you must be kidding.
But you're right, 7z is an under-appreciated format.
The had arc, pak, zip, zoo, warp, lharc, and every Amiga BBS I got on used a different archive format. Everyone had a different opinion on which archive format compressed things in the best way.
I think eventually they decided on lharc when they started to put PD and shareware files on the Internet.
Tar.gz is used because there are instructions for it everywhere and it seems like a majority of free and open source projects archive in it. It is a more popular format than the others right now. Might be because it is an older format and had more ports of it done.
But I really like 7zip, it seems to compress smaller archives, before 7Zip I used to use RAR but WinRAR wasn't open source and 7Zip is so I switched.
With high speed Internet it doesn't seem to matter much anymore unless the file is in over a gigabyte in size. Even then Bit Torrent can be used to download the large files. I think BitTorrent has some sort of compression included with it if I am not mistaken. To compress packets to smaller sizes over the torrent network and then resize them when the client downloads them. That is if compression is turned on and both clients support it.
It happened on DOS too: ZIP, ARJ, RAR, ...
That was back on the days of floppy disks (which usually had at most 1440 KiB) and small hard disks (a few tens of megabytes). Even a few kilobytes could make a huge difference.
As storage and transfer speeds grew, "wasting" a few kilobytes is no longer that much of an issue, and other considerations like compatibility become more important. Furthermore, many new file formats have their own internal compression, so compressing them again gains almost nothing regardless of the compression algorithm.
The reason both ZIP and GZIP became ubiquitous is, IMO, that the compression algorithm both use (DEFLATE) was released as guaranteed to be patent-free, back in a time where IIRC most of the alternatives were either patented or compressed worse. As a consequence, everything that needed a lossless compression method chose DEFLATE (examples: the HTTP protocol, the PNG file format, and so on).
Microsoft ended up adopting LZX for things like CAB and CHM files.
Arch Linux started using lzma2 compression for their packages nearly 6 years ago!
https://www.archlinux.org/news/switching-to-xz-compression-f...
ls /usr/portage/distfiles/ \
| sed 's/.*[.]//g' \
| sort | uniq -c | sort -n -r \
| head -n 6
3377 gz
3051 xz
1656 bz2
295 zip
194 tgz
107 jarThere are others but I can't remember. It's fairly common now.
OSX: tar -xf some.tar.xz (WORKS!)
Linux: tar -xf some.tar.xz (WORKS!)
I had no idea tar could autodetect compression when extracting. (I wonder if this is GNU tar only, or whether the OSX default tar can do it too?) I've been typing `tar zx` or `tar jx` for too long.Bonus: it decompresses to a safely-named subdirectory, but moves the contents of that subdirectory back to the current directory if the archive contained exactly one file. Highly convenient without any risk of accidentally expanding 1000 files into the current directory.
After creating this macro, I've basically never had to care about how to decompress/unarchive anything.
# 'x' for 'eXpand'
alias xx='command atool -x'
# use
% cd $UNPACK_DIR # (optional) (can be the PARENT dir)
% xx foo.zip # or .tar.{gz,xz} or whatever
foo.zip: extracted to `foo' (multiple files in root)
% cd foo/
% ls | wc -l
3
atool actually has many other useful features, but it's worth it just for the extractor.About the only thing that a small would need to be updated at this point is support for a new compressor. (that 2012 release mainly added suppport for plzip)
How does atool deal with cases where there are two versions of the same extractor?
# ~/.atoolrc
path_zip /path/to/preferred/bin/zip
See atool(1) for details. http://linux.die.net/man/1/atoolAs far as I know, different zip formats are not auto-detected. However, it does (optionally) use file(1) to detect the file format, which can be overridden with the 'path_file' option, so a hack may be possible?
The spec doesn't say much on the subject but has this item in the Purpouse section: "Compresses data with a compression ratio comparable to the best currently available general-purpose compression methods and in particular considerably better than the gzip program"
There are times when I do seriously look for the optimum way to do things like this and then there's most of the time I just want to spend brain cycles on more important problems.
I believe that the biggest driver of using old-school ZIP or GZIP is the fact that everyone knows that everything can decompress these formats. And in a modern world of terabyte disks in every laptop, multicore multi-Ghz CPUs, and megabit bandwidth, it isn't worth the effort of using a format that saves an additional 20% on compressed size at the cost of someone not being able to decompress it.
Ian Witten put together the Calgary corpus - https://en.m.wikipedia.org/wiki/Calgary_corpus
> dd if=/dev/urandom bs=1M count=5 | gzip > test
5+0 records in
5+0 records out
5242880 bytes (5.2 MB) copied, 0.438033 s, 12.0 MB/s
> dd if=/dev/urandom bs=1M count=5 | xz > test
5+0 records in
5+0 records out
5242880 bytes (5.2 MB) copied, 1.52744 s, 3.4 MB/s
> dd if=/dev/urandom bs=1M count=5 | bzip2 > test
5+0 records in
5+0 records out
5242880 bytes (5.2 MB) copied, 0.804324 s, 6.5 MB/sIt's not the algorithm per-se that go obsolete, but their use in specific cases until all are diminished. Whether lossy or lossless, eventually other technological advancements renders them unnecessary.
And it seems that strongest algorithm is usually the earliest to be widely adopted; these are almost never toppled.
Just like .gz, look at MP3 or JPEG -- 'better' alternatives exist, but the next widely adopted step will be to eliminate that compression entirely. The first radio station playout systems were hardware MPEG audio compression, and the next most widespread step was to uncompressed WAVs. Even video pipelines based on uncompressed frames are becoming more widespread. Eventually the complexity and unpredictability of compression is shunned for simplicity.
Read the gzip docs and the focus is around compression of text source code, a key use case at the time but barely considered these days -- tar.gz source archives exist almost only out of habit; they could just as well be tar.
For source tar balls though there is basically 0 cost to switching. Since you can download a new compressor in under a minute and should be able to assume that your users are pretty sophisticated. The incentives are similar to media in that cost transfer has to be weighed against the cost of CPE, except that since users supply the CPE the cost is effectively 0 and compression will probably always make sense.
P.S. There are gzip-compatible implementations with tiny per-connection encoding memory requirements (< 1KB).
A few reasons why gzip is still useful to have around:
* Speed is critical for many applications, and so size can take a backseat when performance is critical or resources are low.
* gzip is basically guaranteed to be available everywhere in utility and library forms.
* Download speeds vary and so the faster your pipe, the less the archive size factor will matter, and the faster-worse compression might win out in other comparisons.
* xz doesn't compress every type of data this much better than gzip. I've dealt with scenarios where the difference is consistently less than 2%, and the extra time xz spends is actually a tremendous waste.
Sure, for package downloads where xz files will be significantly smaller it makes sense to save the bandwidth, time and storage space. But it's not 100% cut and dry.
What is the compelling thing here that makes him feel that this is a moral imperative?
$ ncat -C --ssl news.ycombinator.com 443 <<EOF
GET / HTTP/1.1
host: news.ycombinator.com
EOF
Works for me, no compression. Maybe you messed something up or maybe there's a non-compliant proxy between you and the rest of the internet? accept-encoding: identityThat's my problem with GZIP, in this particular use case anyway.
(1) People who still use 56k modems to download content
(2) People who host extremely popular downloads and who want to minimize their outbound bandwidth bills
If you're not one of these two, you almost certainly care more about compatibility and compression time than compression ratio. gzip continues to win on both those fronts, and it explains why it's still the most popular compression format other than ZIP (which is a better choice than gzip if you frequently need to extract a single file from a compressed archive).
Until there's a compression tool released that can compress at wire speed (like gzip) and has a significantly better compression ratio, don't expect the landscape to change much.
> We present a fast compression and decompression technique for natural language texts. The novelties are that (1) decompression of arbitrary portions of the text can be done very efficiently, (2) exact search for words and phrases can be done on the compressed text directly, using any known sequential pattern-matching algorithm, and (3) word-based approximate and extended search can also be done efficiently without any decoding. The compression scheme uses a semistatic word-based model and a Huffman code where the coding alphabet is byte-oriented rather than bit-oriented.
~ $ xx /usr/portage/distfiles/xz-5.0.8.tar.gz
~ $ cd xz-5.0.8/
~/xz-5.0.8 $ ./configure --help | grep -A1 scripts
--disable-scripts do not install the scripts xzdiff, xzgrep, xzless,
xzmore, and their symlinks $ cat zgrep
#!/bin/sh
# zgrep FILE ARGS...
FILE="$1"
shift
gzip -d <"$FILE" | grep "$@"And when packing you can use -a, then tar selects which compression method to use based on the filename.
tar -cv "folder" | xz >"out.tar.xz"
hostA$ tar c mydir | lzop | socat - tcp-listen:1234
hostB$ socat tcp:hostA:1234 | lzop -d | tar x
As the network gets slower and the CPU's get faster, you substitute a more heavyweight compression command.