Lots of interesting imagery, with source included.
My background, I took a bit of liberty with the colors: https://imgur.com/a/OFcsjZn
And an animated version: https://youtu.be/mO7VuqLNK1w
Fun use case to play around with Octrees to try and speed it all up.
JPEG uses a very different technology; it breaks the image into 8x8 blocks, and tries to fit the resulting 64 pixels to a gradient (yes, I know I'm simplifying). So pixels that tend to be "smooth" and "gradient-like" will compress much better than random noise.
I feel it's very much spot on, and true to the math within.
I'll see if I can come up with a clever way do it without having to resort to tricks that move the goalposts (such as having repeated colors).
Okay I think I got pretty close. The png export doesn't have 16.7 million colors, so there must be some error in my logic.
1370 bytes as a human readable plaintext file.
711 bytes as a 7-zip.
658 bytes as a usable compressed svgz file.
159936 bytes as a png.
<?xml version="1.0" encoding="UTF-8"?>
<svg width="256" height="65536" version="1.1" viewBox="0 0 256 65536" xmlns="http://www.w3.org/2000/svg" xmlns:cc="http://creativecommons.org/ns#" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#">
<defs>
<linearGradient id="a" x2="256" y1="32768" y2="32768" gradientUnits="userSpaceOnUse">
<stop stop-color="#f00" offset="0"/>
<stop stop-color="#ff0" offset=".166666667"/>
<stop stop-color="#0f0" offset=".333333333"/>
<stop stop-color="#0ff" offset=".5"/>
<stop stop-color="#00f" offset=".666666667"/>
<stop stop-color="#f0f" offset=".833333333"/>
<stop stop-color="#f00" offset="1"/>
</linearGradient>
<linearGradient id="b" x1="128" x2="128" y1="0" y2="65536" gradientUnits="userSpaceOnUse">
<stop offset="0"/>
<stop stop-color="#808080" stop-opacity="0" offset=".5"/>
<stop stop-color="#fff" offset="1"/>
</linearGradient>
</defs>
<metadata>
<rdf:RDF>
<cc:Work rdf:about="">
<dc:format>image/svg+xml</dc:format>
<dc:type rdf:resource="http://purl.org/dc/dcmitype/StillImage"/>
</cc:Work>
</rdf:RDF>
</metadata>
<rect width="256" height="65536" fill="url(#a)" stroke-width="24" style="mix-blend-mode:normal"/>
<rect width="256" height="65536" fill="url(#b)" stroke-width="24" style="mix-blend-mode:normal"/>
</svg> <svg width="4096" height="16384" version="1.1" viewBox="0 0 4096 16384" xmlns="http://www.w3.org/2000/svg" xmlns:xlink="http://www.w3.org/1999/xlink">
<defs>
<linearGradient id="a" x2="4096" y1="8192" y2="8192" gradientUnits="userSpaceOnUse">
<stop stop-color="#f00" offset="0"/>
<stop stop-color="#ff0" offset=".125"/>
<stop stop-color="#0f0" offset=".25"/>
<stop stop-color="#0ff" offset=".375"/>
<stop stop-color="#00f" offset=".5"/>
<stop stop-color="#f0f" offset=".625"/>
<stop stop-color="#f00" offset=".75"/>
<stop stop-color="#000" offset=".875"/>
<stop stop-color="#fff" offset="1"/>
</linearGradient>
<linearGradient id="b" x1="2048" x2="2048" y2="16384" gradientUnits="userSpaceOnUse" xlink:href="#a">
<stop stop-color="#808080" offset="0"/>
<stop stop-color="#808080" stop-opacity="0" offset=".25"/>
<stop stop-color="#fff" offset=".5"/>
<stop stop-color="#fff" stop-opacity="0" offset=".75"/>
<stop offset="1"/>
</linearGradient>
</defs>
<rect width="4096" height="16384" fill="url(#a)" stroke-width="24" style="mix-blend-mode:normal"/>
<rect width="4096" height="16384" fill="url(#b)" stroke-width="24" style="mix-blend-mode:normal"/>
</svg>An image with the same pixels but randomly arranged cannot really be compressed.
When you save in image in JPG, it's compressed using an algorithm that gets it "pretty close" to the source image. The software doing the JPG decoding (browser, image viewer, etc.) basically reverses this algorithm to display the image back to you. For JPG, this compression is "lossy" and so you've lost detail from the original source image.
PNG is a lossless image format, but it basically works the same way without sacrificing the source image's quality.
The "standard" for an image format dictates how an encoder creates the image and how a decoder displays the image. Since everyone is "on the same page," the individual image files only need to contain what they need to -- basically, "this pixel = RGB(0,1,2)".
Hope this makes sense and helps.
Gotta ask. Has anyone worked on image compression using machine learning? (that sounds like an obvious thing to do). It would be funny to end up with an algorithm no one understands.
Given both the financial value of image compression (given the amount of video shoved down the net) and the asymmetry of codec (it's OK to use resources to compress, not so much for decompress), I'd expect some real money to be spent in this area.
Dunno about the state of the art, but pretty much every ml tutorial has a section titled "Image Compression Using Autoencoders" right at the beginning. It's the Hello World of ml. The perceptual quality vs file size curve for such a simple network is pretty mediocre, but I'm sure you could do better if that was your goal.
Tangentially related is the crash bandicoot game on the playstation. The developers figured out that untextured polygons were way faster to draw than textured polygons, and so they made the player character model out of tons of tiny colored polygons rather than fewer larger textured polygons. The result was a significantly better looking graphic for the same rendering time.
Now your question is basically reduced to "how can text files with same number of bytes, each having ALL the ascii codes, compress to files of such different size". The answer is that it MUST necessarily be so. You can't have a one-to-one map from [2]^N to [2]^K where K < N.
Except that you've already noted that JPG is a lossy compression scheme, so it doesn't matter that a one-to-one map isn't possible.
For the ordered image: "Make a 16 by 16 grid of boxes, each getting redder from left to right and top to down. Inside each box make a 16x16 grid of boxes getting greener, and inside each of those boxes make a 16x16 grid of pixels getting bluer". Done.
Vs. the random image: "Make a blueish red pixel with a bit of green. Then make a reddish blue pixel with a a moderate amount of green. Then make a brownish pixel. Then make a greenish pixel..." That will go on for about 16 million sentences!
Image compression is simply coming up with a language that's good for describing images, and then an algorithm for writing succinct descriptions in that language. Rather than being human friendly languages (a modest number of complex words made from an alphabet), they are computer friendly languages (a ton of simple words made from 0s and 1s.)
JPEG compression separates out the image into "component" images. One component is brightness, another is hue, and a third is saturation. Each of those components map nicely to how human vision works. In particular, brightness needs high resolution and precision. Hue needs high precision but resolution doesn't matter. Saturation doesn't need high resolution or precision. Therefore, each component can be compressed differently and independently from each other. The component images are further broken up into a grid of 8x8 boxes, and then each of those boxes are approximated via a weighted sum of reference box images (the encoder and decoder have a dictionary of reference boxes they both agree to use.) For each box, only the weights are saved. The weights themselves can have varying precision, and that's basically you're controlling when you set the "jpeg quality". Higher quality jpegs have more precision in their weights, and lower quality jpegs have less precision in their weights.
PNG works like this. You can run each horizontal row through any one of a variety of delta encoders that are suitable for different situations. The goal is to minimize the range of values and maximize the repetition before you pass the encoded deltas through a dictionary based compressor. Pictures like this are near optimal for this approach.
It uses simulated annealing to arrange the colours smoothly.