Show HN: Neural Image Compression Demo
colab.research.google.com
colab.research.google.com
This project is based on the paper "High-Fidelity Image Compression" by Mentzer et. al. [3] - this was one of the most interesting papers I've read this year! The model is capable of compressing images of arbitrary size and resolution to bitrates competitive with state-of-the-art compression methods while maintaining a very high perceptual quality. At a high-level, the model jointly trains an autoencoding architecture together with a GAN-like component to encourage faithful reconstructions, combined with a hierarchical probability model to perform the entropy coding.
What's interesting is that the model avoids compression artifacts associated with standard image codecs by subsampling high-frequency detail in the image while preserving the global features of the image very well - for example, the model learns to sacrifice faithful reconstruction of e.g. faces and writing and use these 'bits' in other places to keep the overall bitrate low.
The overall model is around 700MB - so transmitting the model wouldn't be particularly feasible, and the idea is that both the sender and receiver have access to the model, and can transmit the compressed messages between themselves.
If you have any questions or notice something weird I'd be more than happy to address them.
---
[1] Colab Demo: https://colab.research.google.com/github/Justin-Tan/high-fid...
[2]: Github: https://github.com/Justin-Tan/high-fidelity-generative-compr...
[3]: Original paper: https://hific.github.io/
[4]: Sample reconstructions: https://github.com/Justin-Tan/high-fidelity-generative-compr...
One interesting failure model is that images dominated by high-frequency detail require a relatively large bitrate to store - see e.g. the last example in the Github README with the weird brickwork. Even though the model was trained to produce compressed representations with a soft constraint on the maximum bitrate, the filesize of the representation for this particular image is something like 60% above the nominal maximum.
The model in the demo is a lossy compression method because it first projects the input to a lower dimensional space and performs quantization of this representation to integer values so the result can be ultimately entropy coded. It uses the mean-scale hyperprior model introduced in [1] to estimate the necessary probability distributions in the lower-dimensional space for entropy coding.
[1]: https://arxiv.org/abs/1811.12817 [2]: https://arxiv.org/abs/1802.01436
From the README
> The generator is trained to achieve realistic and not exact reconstruction. It may synthesize certain portions of a given image to remove artifacts associated with lossy compression. Therefore, in theory images which are compressed and decoded may be arbitrarily different from the input. This precludes usage for sensitive applications. An important caveat from the authors is reproduced here:
> "Therefore, we emphasize that our method is not suitable for sensitive image contents, such as, e.g., storing medical images, or important documents."
As an example of this going wrong previously, xerox had once implemented compression based on deduplicating duplicate parts of documents. Obviously numbers contains tons of duplicate symbols (digits). The problem was that the scanner software deduplicated different numbers with each other, leading to wrong numbers.
http://www.dkriesel.com/en/blog/2013/0802_xerox-workcentres_...
However, the model does learn a conditional probability distribution over a lower-dimensional representation of the original image - this is unavoidable as entropy coding requires a distribution over discrete symbols. The GAN is almost auxiliary and not a central component of the model - in fact, you can get very good results without the GAN, but does seem to result in visually superior reconstructions.
Yes, I understand this is a lossy compression method - what I was proposing is to have the compressor as a final pass take the predicted output image, and subtract it from the original pixels. This gives you a delta between the predicted image and the original image. You can then compress that delta losslessly, and store it alongside the output of this model - if the predicted image is close enough to the original image then you've significantly reduced the amount of entropy in the delta, making it highly compressible.
This is how some domain-specific lossless compression algorithms work, e.g. DTS-HD Master Audio
The free tier for traffic is "15GB of Data Transfer Out", after that it's 9 cents per GB. https://aws.amazon.com/s3/pricing/?nc1=h_ls (under "Data transfer"). Check your AWS bill!
I assume most people just want to see the images, forcing them to recompute them is a waste of resources. Even just storing a version of the colab with the results present would help a lot.
> [4]: Sample reconstructions
The text in the reconstructed image in the third row looks different, the word phonomat is quite garbled, information looks a bit funny.
Also check out the previous discussion on HN:
One shortcoming is that this current model is non-adaptive - which means that the target rate is fixed. So to achieve different target compression rates you would have to train multiple models in different rate regimes. In the Colab demo there is the option to select between 3 different models trained with a target bits-per-pixel (bpp) rate at 0.14bpp, 0.30bpp, and 0.45bpp, respectively - higher rates correspond to more higher-fidelity reconstructions, at the expense of a lower compression ratio. The default is the `HiFIC-med` model (and this is what the all samples in the README were generated with), but the model trained at the highest bitrate should have less obvious imperfections.
There's also an aspect to the distortion that can be attributed to the entropy coding process rather than the model itself - currently the system clips values outside a certain probability range, resulting in artificial distortion - a fix is in the pipeline though.
(I just use the exmaple image that is already in the notebook)
Original: https://i.imgur.com/Q66mHTD.png Result: https://i.imgur.com/4R6qn8e.png
There are lots of random spots on the image, and the brightness level changes totally.
Sure, 5232 kB to 124 kB is impressive, but people would probably prefer a badly compressed JPEG over this, since at least JPEG artifact is predictable (and if image isn't displayed in 100%, the artifact would be less obvious, unlike brightness change and spots in this result).
Edit: I just saw the result in https://hific.github.io/ for the same picture, but that one has none of these flaws (no brightness change, no weird spots here and there) with even smaller filesize. Why?
As for the random spots, that's an artifact of the entropy coding algorithm. In principle this is lossless but there is some distortion because I'm using a custom vectorized version of an rANS encoder and it's hard to encode overflow values in a vectorized fashion, I'm working on this though. If you can live with really slow decoding times (2-3mins) then you can disable vectorization to eliminate these small imperfections entirely.
As for the comparison to the official model, that's mainly because of compute constraints v. Google (this is just my weekend project). My model uses a smaller architecture and was trained for only 4e5 steps versus the 2e6 steps they reported in the paper - even then it took 4+ days on AWS! The model is also trained on the Openimages dataset, which is presumably much smaller and more noisy than the massive internal dataset Google used.
[1] https://colab.research.google.com/github/Justin-Tan/high-fid...
Edit: original is on the right
It's incredibly hard to change the default file format on the web, but there's an opportunity to switch libjpeg to a decoder with much more realistic output images.
otherwise seemed to work
# Setup model
I get an error in the function call 'prepare_model'
UnpicklingError: invalid load key, '<'.
One solution is to download (and upload to Colab) the models manually in /content/checkpoint/
``` # Setup model
I get an error in the function call 'prepare_model'
UnpicklingError: invalid load key, '<'. ```
Try rerunning the download cell if you experience this - the models downloaded should be around 1.5-2GB, so if the checkpoints are 100kB in size, the download's gone wrong.
That was the error I got!