Consider images of slides for a presentation that are text on a flat background. If you know the value of the pixel just to the left of the current pixel, then if you guess that the current pixel will be the same, you will be right most of the time. This is obviously not true for random noise. Consider a really simple compression scheme where a pixel is stored as a single bit of a 1 if it is the same color as the previous pixel, and the color is stored as usual, but with an additional 0 bit prepended to the color. When you guess wrong, you pay a tax of 1 bit, but when you guess right you save N-1 bits where N is the number of bits per pixel.
For random noise, this will grow the input quite a bit, but for simple flat-shaded graphics it will shrink the input quite a bit.
We live in a world where it is common for blue sky to occupy a portion of the frame and green grass to occupy another portion of the frame. Since images captured of our world exhibit repetition and patterns, there are opportunities for lossless compression that focuses on serializing deviations from the patterns.
If that's the case, then not only could we rely on the results of everything we've decompressed so far to use for prediction (which is like one-sided image in-painting), but we also could store a few bits of semantic information (e.g. from an image-net-based CNN, from face detection) about the content of the original image before re-compression, and use that semantic information for prediction as well via some generative model. All of this would obviously be trading computation for storage/bandwidth, but it this seems like an exciting direction to me. Again, nice work.
As for having the mega-model that predicts all images better: well it turns out with the lepton model out you only lose a few tenths of a percent by training the model from scratch on each images individually. We have a test case for training a global model in the archive (it's https://github.com/dropbox/lepton/blob/master/src/lepton/tes... ) That trains the "perfect" lepton model on the current image then uses that same model to compress the image (It's not meant to be a fair test, but it gives us a best-case scenario for potential gains from a model that has been trained from a lot of images) and in this case it doesn't gain much, even in a controlled situation like the test suite.
However the idea you mention here may still be a good idea for a hypothetical model--but we haven't identified that model yet.
I can't read the article due to technical constraints, but understand that e.g. JPEG has a lossy quantisation pass followed by a lossless encoding/compression pass of the result of the first stage. If they're reproducing bit-identical result to the input JPEG, it must be a (very good) optimisation of the latter stage. [How'd I do?]
That would actually compress rather well. A hard to compress random image would not look like TV snow (white dots, with space between them), but rather randomly colored dots that are continuous in the image.
https://dsp.stackexchange.com/questions/2010/what-is-the-lea...
In this case they are making use of the fact that JPEGs store a certain mathematical formulation of a picture. It turns out JPEG doesn't store that mathematical formulation very well, so you can squash it loselessly into a better formulation, then later turn it back into the original JPEG.