Finally Understanding PNG (2017)
compress-or-die.com
compress-or-die.com
It’s a bit of a shortcut: in reality PNG uses DEFLATE, and it’s DEFLATE which has a 32kB window.
And it’s not exactly performance reasons but memory budget, especially as deflate was designed for streaming, both when compressing and decompressing: when decompressing you need a buffer at least as large as the window, and when compressing you need that plus most likely an index of some sort (so you don’t have to look for matches linearly through the buffer).
This is more relevant with the context that Phil Katz created deflate in 1991, high-end desktops reached high single-digit megabytes of RAM but the low end might not even be there (Windows 3.0 required 1MB to run decently).
By comparison, if memory serves zstd uses a window of 8MB (or 16 depending on setting), though configurable anywhere from 1k to 512MB
For reference: https://en.wikipedia.org/wiki/X86_memory_segmentation#Real_m...
HDD:
https://en.wikipedia.org/wiki/Hard_disk_drive
Diskettes:
https://en.wikipedia.org/wiki/Floppy_disk#3+1%E2%81%842-inch...
I'm commenting "HDD" as written in: "1.44 MB. This is also pretty much the size of an uncompressed BMP image (and incidentally the size of a 3.5 " HDD ... for all who can still remember" in the current text of the article.
It was a floppy disk, not a hard disk. The magnetic surface in an FDD disk is floppy the magnetic surface in a HDD is hard (rigid).
Maybe I'm wrong but I was under the impression that a "floppy" was called that because the original ones - 5.25", were floppy.
Nitpick: The first floppy disks were 8" in size.
I imagine there won't be a mainstream culture war over it like there is with GIF and JIF though.
I wonder how much bandwidth I've wasted for people over the years because I've saved poorly optimised versions from Photoshop. Are there automated tools that can be included in a CI pipeline to help with this?
[1] https://packages.debian.org/unstable/optipng
[2] https://packages.debian.org/unstable/pngcrush
[3] http://www.olegkikin.com/png_optimizers/Which contains all of those and just tries them all. Been meaning to pop it all out and make a bash script of them to automate some of my stuff on a cron job instead of using a GUI.
I recall PNG8 not being well supported in some applications. I've run into this with game engines, which I found strange since asset size is an important thing to keep in mind.
There were some IE-specific hacks you could use with their proprietary CSS extensions but they were really annoying. I think IE7 finally fixed the transparency bugs but since there was still a huge IE6 cohort it was still a crapshoot relying on transparent PNGs.
One particular trick is getting images with visually low number of distinct colors, like logos and screenshots, through pngquant. You don't have to worry whether the actual number of colors is less than 256, how to posterize or filter the image to decrease that number, or how to suppress artificial colors from font anti-aliasing — the tool measures the visual likeness, and provides the means to set the acceptable limit so it would leave unsuitable images intact. Obviously, when you're dealing with some long-term archival storage, or assets for further editing like stock clipart, such presentation-oriented lossy optimization should not be used.
Other automatic tools are referenced in the article, though you should treat them as a nice sanity check that prevents images holding raw uncompressed data or simple two color graphs saved as 4 byte RGBA from appearing in your project, and not a magical solution that would send it to the sky performance-wise.
Bandwidth (and loading time) is usually wasted on complex high resolution images like photos or combined graphics with photographic details. These can't really be optimized well enough in PNG with naive lossless or lossy methods, and have to be converted to JPEG to reach sane size or regenerated as independent layers that use different compression.
Not really. I found that for many many many pictures, not using a filter turns out to be the smallest option. It's really just that DEFLATE is much better than LZW as used by GIF.
Though to be honest using Palette for colors tends to produce a huge saving for a lot of image types (those with 256 or fewer colors, so documents, not natural scenes) making filter selection a small added bonus.
[0]: https://github.com/EliotJones/BigGustave/blob/master/src/Big...
That paragraph makes it sound like this article was written two decades ago. pngcrush, which does bruteforce, still takes only a minute or two on average (less than screen-sized) files.
Instead, pngcrush is using yet more heuristics to guess what else is likely to work and then trying some of the other options that wouldn't be picked by the defaults.
By default it only tries a few approaches that it deems most likely to work, but even with the "brute" parameter it does not truly try everything, again, that would take too long, and pngcrush "brute" is already very slow, it just tries all the methods it would ever have used on any image.
Actual brute force of a relatively tiny 32x32 image would require trying all five filters for each line in combination with each such filter for each other line or else you might miss a combination which, astonishingly, was smaller even though it wasn't immediately obvious why. I think that's 5 raised to the power 32 combinations.
This could be trivially extended to multiple lines of look-ahead, though at exponential memory and CPU cost.
Probably dog slow though.
The problem is, while you know which outputs compressed best so far you can't be sure which will compress best in combination with future outputs you haven't as yet filtered.
Imagine for example that one filter for the current row gives what seems like a random sequence that doesn't compress very well. Your proposed scheme rejects this filter, it compressed poorly, so don't use it. But, what if it turns out that every future row can be compressed to the exact same sequence with some other filter? Deflate can compress the rows extremely well together, even though the first one on its own did not compress.
Would be a fun weekend project...
In your example, even if the first row is more poorly compressed with one filter, could future rows "catch up" the loss? Unless you have a specific counterexample, I don't think it can, because the filtration operates before compression and there's no feedback as such.
The compressor consumes the entire filtered data. You may have been thinking it goes something like
Compress(Filter(row1,method=2)).Compress(Filter(row2,method=4)).Compress(Filter(row3,method=1)) and so on.
But that would be terribly wasteful. What they do is actually closer to
Compress(Filter(row1,method=2).Filter(row2,method=4).Filter(row3,method=1) and so on
Does that make it clearer that if the results of two Filter steps are more or less similar, even if those results compress badly on their own, it helps them compress better when they're both nearby ?
Huffman was a Doctor of Science.
Also, (2017).
Edit: Couldn't OP's email. I don't have Twitter.