pigz: A parallel implementation of gzip for multi-core machines
github.com
github.com
One trick I discovered is that tools like pigz can be used to both accelerate the compression step and also copy to cloud storage in parallel! E.g.:
pigz input.fastq -c | azcopy copy --from-to PipeBlob "https://myaccountname.blob.core.windows.net/inputs/input.fastq.gz?..."
There is a similar pipeline available for s3cmd as well with the same benefit of overlapping the compression and the copy.However, if your tools support zstd, then it's more efficient to use that instead. Try the "zstd -T0" option or the "pzstd" tool for even higher throughputs but with same minor caveats.
PS: In case anyone here is working on the above tools, I have a small request! What would be awesome is to automatically tune the compression ratio to match the available output bandwidth. With the '-c' output option, this is easy: just keep increasing the compression level by one notch whenever the output buffer is full, and reduce it by one level whenever the output buffer is empty. This will automatically tune the system to get the maximum total throughput given the available CPU performance and network bandwidth.
--adapt[=min=#,max=#]
zstd will dynamically adapt compression level to perceived I/O conditions. Compression level adaptation can be observed live by using command -v. Adaptation can be constrained between supplied
min and max levels. The feature works when combined with multi-threading and --long mode. It does not work with --single-thread. It sets window size to 8 MB by default (can be changed manu‐
ally, see wlog). Due to the chaotic nature of dynamic adaptation, compressed result is not reproducible.Similarly, on a local speed test (SSD -> SSD), using a fixed compression level was much faster than --adapt.
"" note : at the time of this writing, --adapt can remain stuck at low speed when combined with multiple worker threads (>=2). ""
There are some ADVANCED COMPRESSION OPTIONS --zstd tunables that might help.Leave wlog alone unless you're willing to store the value out of band and pass it in again during decompression.
hashLog, bigger number uses more memory to compress but is often faster.
chainLog smaller number compresses faster, but worse ratio.
In your use case monitoring general system utilization to identify bottlenecks might also help. My gut instinct is that you might already have hit a memory bandwidth limit for the platform, at which point REDUCING the hashLog until it fits within your intended performance budget might yield better bandwidth results. Reducing the chainLog value might have the same effect.
[1] https://en.wikipedia.org/wiki/TCP_congestion_control#TCP_BBR
I believe this is because it puts similar strings together in the tree gzip uses. clumpify.sh from bbtools is one example.
https://jgi.doe.gov/data-and-tools/software-tools/bbtools/bb...
You may also filter out duplicates and depending on what you do correct base errors or get rid of low complexity reads.
The bgzf format was invented for bioinformatics, and BAM files are bgzipped by default.
It blew people's minds when we couple implement huge projects that they could never get of the ground for years in just a matter of months because of 'this one cool trick'.
Man I miss Santa Cruz ಥ﹏ಥ worst mistake of my life to leave there.
Why not use some binary format like BSON? Compressing with gzip works but then you can't query it compressed
And text is easier to work with in any hacked up script.
There are some binary formats (BAM, I think), but people often prefer the text format anyway. When compressed, the size is pretty much the same
They do, but you need to uncompress to read the N-th letter
https://github.com/madler/pigz/blob/master/try.h
https://github.com/madler/pigz/blob/master/try.c
which implements try/catch for C99.
[1] https://kotlinlang.org/docs/exceptions.html#java-interoperab...
Of course you'd have to deal with exceptions since you're running on the JVM and want to Interop with Java code, that doesn't mean it's idiomatic code.
All three of those languages actually have exceptions, they just don't encourage catching exceptions as a normal way of error handling
Also, while the trend now seems to be for newer languages to encourage use of things like result types, one of the main reasons for that is that in current languages it is easier to show that functions can potentially failure in the type system using result types rather than exceptions.
Otherwise, there isn't necessarily inherently a strong reason to prefer one or the other, and it's possible that future languages will go back to exceptions but have a way to express that in the type system using effects, etc.
[1] https://github.com/postgres/postgres/blob/master/src/include...
*edit, as ac29 mentions below, just use zstdmt. In my quick testing it is approximately 8x faster than pbzip2 and gives better compression ratios. Wall clock time went from 41s to 3.5s for a 3.6GB tar of source, pdfs and images AND the resulting file was smaller.
megs
3781 test.tar
3041 test.tar.zstd (default compression 3, 3.5s)
3170 test.tar.bz2 (default compression, 8 threads, 40s)For everything else, there's zstd (also natively multithread)
I generally use lzip for data that is important to me.
parallel-friendly, trades off compression level for speed
I recall learning the technique from Tumblr's eng blog.
https://engineering.tumblr.com/post/7658008285/efficiently-c...
Separately, at Tumblr I vaguely remember examining some alternative to pigz that was consistently faster at the time (11 years ago) because pigz couldn't parallelize decompression. Can't quite remember the name of the alternative, but it had licensing restrictions which made it less attractive than pigz.
Edit: the old fast alternative I was thinking of is qpress, formerly hosted at http://www.quicklz.com/ but that's no longer online. Googling it now, there are some mirrors and also looks like Percona tools used/bundled it. Not sure if they still do or if they've since switched to zstd.
SQL backups were simply a bash script using Pigz, running on a cron job. Simple times!
This is what pigz is doing: shooting the input into blocks, spreading the compression of these blocks over different threads so multiple cores can be used, then joining the results together in the right order.
It is the very same property of the format that gzip's own --rsyncable option makes use of to stop small changes forcing a full file send when rsync (or similar) is used to transfer updated files.
The idea is as simple as it is clever, one of those "why did I not think about that?" ideas that are obvious once someone else has thought of it, so adds little or no extra risk. A vulnerability that uses gzip (a "compression bomb") or can cause a gzip tool to errantly run arbitrary code, is no more likely to affect pigz than it is the standard gzip builds.
There are multiple implementations of zlib that are faster than the one that ships with GNU gzip, and yet they haven't been incorporated.
There are also just better algorithms if compatibility with gzip isn't needed. zstd, for example, supports parallel compression, and is both faster and compresses better than gzip.
I suspect to keep the standard gzip as simple, small, and stable, as possible. It does the job, does it well enough, with minimal dependencies, has done for many years, and can do so in a wide array of systems including very small environments (in part due to the minimal dependencies).
Core tools like that typically don't get major updates, just security & stability patches as needed and maybe the occasional safe & easy change for performance reasons or to widen the number of supported environments.
Given almost all desktops/servers/mobile phones are likely to be multi core these days, if gzip gained multithreading on desktop it could save time and energy for the whole planet potentially, that seems like a worthwhile benefit?
I think there is a case for including it as a selectable option, as with --rsyncable, unless this adds extra dependencies (pthreads was mentioned in other comments).
I try to avoid it generally in my coffee too, but for something that could potentially offer a measurable benefit in a world of multi core-solid state computing, would it not be "worth it"?
That’d be a pretty funny compression algorithm. You listen to a .mpfoo file, and you’ll hear the whole song, we promise!
zstandard is faster and slightly better compression at speed selection settings that are equivalent to gzip, in addition to having the ability to compress stuff at a much greater ratio, optionally, if you allow it to take more time and cpu resources.
https://gregoryszorc.com/blog/2017/03/07/better-compression-...
It works because gzip streams can be tracked together as a single stream, at the start of each block is an instruction to reset the compression dictionary as if it is the start of a file/stream (which in practise it is) so you just have to concatenate the parts coming out of the parallel threads in the right order. These resets cause a small drop in overall compression rates but this is small and can be minimised by using large enough blocks.
HAVE_ZLIB : zstd can compress and decompress files in .gz format. This is ordered through command --format=gzip. Alternatively, symlinks named gzip or gunzip will mimic intended behavior. .gz support is automatically enabled when zlib library is detected at build time. It's possible to disable .gz support, by setting HAVE_ZLIB=0. Example : make zstd HAVE_ZLIB=0 It's also possible to force compilation with zlib support, using HAVE_ZLIB=1. In which case, linking stage will fail if zlib library cannot be found. This is useful to prevent silent feature disabling.
zstandard's own zstandard-own-compression format is its own separate much newer thing.
zlib, for instance
(unprivileged commands follow)
dd if=/dev/zero of=~/zeros bs=1M; sync; rm ~/zeros
Compressing on the fly can be slower than your network bandwidth depending on your network speed, your processor(s) speed, and the compression level, so you typically tune the compression level (because the other two variables are not so easy to change). Example backup:
(privileged commands follow)
pv < /dev/sda | pigz -9 | ssh user@remote.system dd of=compressed.sda.gz bs=1M
(Note that on slower systems the ssh encryption can also slow things down.)
Some sharp people may notice that it's not necessarily a good idea to back up a live system this way because the filesystem is changing while the system runs. It's usually just fine on an unloaded system that uses a journaling filesystem.
The zerofree command looks useful, but I don't know how portable it is. The dd method works across many platforms (such as AIX).
Because compression programs are as high-hanging fruit as you can get, and parallelizing them can only be done once.
You won't see any improvement from parallelization on this type of data.
A decade ago I implemented parallel decompresssion for pigz. This is used in Solaris kernel zone suspend and resume, which was the reason I did the work. I submitted a PR for it but madler never got around to reviewing and merging it. Since then there has been a lot of code churn that makes it a pain to apply to the current version.
So interesting if they implemented and tested that and get only a marginal CRC speedup.
edit: someone here seems to observe a ~ doubling with unpigz vs zcat: https://unix.stackexchange.com/a/363739
https://www.linuxjournal.com/content/parallel-shells-xargs-u...
https://web.archive.org/web/20160313033123/https://blogs.ora...
[1] https://docs.oracle.com/cd/E88353_01/html/E37839/pigz-1.html
tar --use-compress-program="pigz -0" ...
Tar is a standalone archive format, you can create tar archives with the tar command and not use any compression utility at all. Here is an example that creates an uncompressed tar archive directly with no compression:
tar -cf $directory.tar $directory
You can also pipe the created archive to stdout instead of to a file if you want to: tar -c $directory | wc -c
# send to a remote system
tar -c $directory | ssh dd of=~/$directory.tar
You can actually compress the archive by just piping it to a compression command: tar -c $directory | zstd --stdout > $directory.tar.zst
The above pipeline is probably very similar to what tar is doing internally when you pass "--use-compression-program".So in your case using "pigz -0" is totally useless since tar creates an uncompressed archive by default. You can totally omit the "--use-compression-command" flag to do what you want.
By "after", I don't mean the whole archive is created before the compression part runs, but rather that tar can stream the the archive to stdout as it's being created. Tar needs to be able to do this since it was intended for tape drives, and seeking around on a tape drive is extremely slow.
It's actually pretty interesting that we still use tar so widely even though tape drives are not in common use! A new archive format that supports random access during reads (and writes maybe?) would be pretty interesting, but tar works pretty well so there isn't that much of a reason to create alternatives.
FWIW, I configure BackupPC to use pigz instead of gzip without any issues.