170 karma · joined January 7, 2016
Additionally, in pipelines, you would only decompress your image once and compress it once. Usually, when people build pipelines based on individual tools passing png's to each other, these steps are repeated n times for n elements of the pipeline.
bzip2 < giantimage.ff.bz2 | header-checker
Inside the header checker, you just read in the first 16 Bytes ("farbfeld" + width + height), do whatever you like with the data and exit. bzip2 would then receive an SIGPIPE. Up until this point, it would have probably filled the pipe buffer (~16K) with data and then blocked.
Receiving SIGPIPE, I'm certain bzip2(1) will cancel any further decompression attempts.
You see, you could even carry on with your image-processing in case you like the header data and want to proceed. The nice pipe-system already handles that for you. :)
how come people assume there are only Little Endian machines. This mmap-technique wouldn't work on Big Endian like this, this is inaccetable!
However, you can still mmap and properly call ntohs() on each color channel value (if you access it). These functions won't hurt performance too much anyway, if at all. If you show me a measurable difference between using LE and BE, I owe you a beer, okay? :)
on a technical side, lzma is a better choice. I might add a lzma recommendation, however, in soma cases lzma performs worse than bzip2.
also, decompression speeds matter, not compression speeds.
in the end, use what you prefer. the compression is not mandated by the spec. the main use for farbfeld is also rather thought to be as a piping format. it could also be used for storage, but that's not the main point.
(2^64)^((2^32-1)(2^32-1)) ~ 10^(10^21)