Massive Speed Gains via Parallelized BZIP2 Compression
hackercodex.com
hackercodex.com
Er, not really. How about...
"pbzip2 reduced running time by 80%."
"pbzip2 took only 20% as long as bzip2 did."
"pbzip2 is five times faster."
It doesn't require use of large temporary files like pxz, and the xz files it produces can also be decompressed in parallel.
Can I suggest adding installation packages for: brew ubuntu (at least through launchpad)
The only thing I wish for is more documentation and inclusion in xz-utils (which I know is not up to you, but I'm hopeful).
The author of xz is working on a parallel compression implementation of his own, which will hopefully be in future versions of xz.
Edit: And in this case, bz2 even gave better compression ratios if I recall correctly.
Why not use parallelized versions of xz such as pxz or pixz? Because pre-built pxz/pixz packages are nearly non-existent. When that changes, I'll consider switching formats.
Also, note that decompressing bzip2 is very slow, xz usually beats it by a factor of two or more.
The fact remains, however, that I can't seem to find a simple way to install a parallelized version of xz. Perhaps I'll post an issue in the Github issue tracker for pixz and see if we can't resolve that. :)
It would be even better if files are arranged in such way that contentwise they are similar one after each other.
I'm not compression expert, just average user since the early days of what we used to call it back in the day "solid" compression.
It doesn't seem unlikely that that's what they're doing, considering you can't pipe data to it on stdin.
Reimplementation of things like compression algorithms seems very math/algorithm-heavy and thus amenable to functional programming. How about the Haskell/OCaml guys re-implement a bunch of Un^x style utilities for us?
I've used it to great effect during incident response when I needed to search through hundreds of gigs of logs at a time.
File cache effect? He should cold reboot first (not sure how you force the file cache out on OSX/linux, on Windows I do it with SysInternals RamMap) and try in different order.
It could still be faster, but he could really be measuring I/O that was done in the first case, and not in the second.
It's also strange that .tar files are used, not tar.bz2 or .tbz (if such extension makes sense)
echo 1 > /proc/sys/vm/drop_caches
http://linux-mm.org/Drop_CachesNot sure why you're surprised that I used .tar files for the compression testing. As I mentioned in the article, most of the time I'm creating bzipped tarballs from directories of files, so it made sense to use what is, for me, a common real-world use case. Your mention of tar.bz/.tbz makes me think there's some misunderstanding, since clearly I wouldn't want to test compression of already-compressed files. But perhaps it's I who am misunderstanding your suggestion. Please feel free to enlighten me. :)
lbzip2 is able to decompress single streams using multiple threads, which apparently pbzip2 cannot do. See the thread beginning with http://lists.debian.org/debian-mentors/2009/02/msg00098.html
That said, I guess using the "alternatives" framework in Linux it would be reasonably easy (and transparent) to support the parallel version of each tool as replacement to the regular one.