HNHacker News
TopNewBestAskShowJobs

tonic_section

81 karma · joined March 23, 2018

https://github.com/Justin-Tan

https://justin-tan.github.io/

submissionscomments
tonic_section··on State-of-the-Art Image Generation Models
You mixed up implicit and explicit models. For anyone interested in the difference - implicit models such as GANs don't allow you to evaluate the probability density over datapoints - you can only sample from some surrogate model of the distribution learned by minimizing some 'distance' between the surrogate and the true empirical distribution.

'Explicit' models (I think this term is nonstandard) parameterize the density directly and modify the parameters via maximum likelihood. This allows one in theory to both directly evaluate the density and sample from the learned distribution. VAEs (only give a lower bound on the density), autoregressive models, and normalizing flows all fall under this category.

Note that while it is theoretically possible for 'explicit' models to go in both directions (sample and evaluate), one direction may be much more efficient than the other for certain models. e.g. for autoregressive models you can read the first two pages of [1] for a good explanation of why.

[1]: https://arxiv.org/abs/2102.11495

tonic_section··on Differentiable Dithering
The objective function used in these lossy neural compression schemes usually takes the form of a rate-distortion Lagrangian - the rate term captures the expected length of the message needed to transmit the compressed information and the distortion term measures the reconstruction error. So it wouldn't be able to cheat like in your example, because this would incur a high value of the loss through the rate term.
tonic_section··on Differentiable Dithering
Usually the lossless encoding is offloaded to a standard entropy coder, e.g. arithmetic, ANS, etc. because these approach the theoretical minimum rate given by the source entropy pretty closely, so there wouldn't be a point building a fancy differentiable replacement.
tonic_section··on Differentiable Dithering
There are a couple of solutions which work empirically - as you mentioned, one solution is a dithering-like differentiable relaxation where uniform noise is added, which simulates quantization, or just to ignore the quantization operation when taking gradients, essentially treating it as an identity operation in the backward pass.
tonic_section··on Differentiable Dithering
In terms of the decoded image, yes - it's very unlikely you would get something substantially different from the original image. But in terms of the bitrate it's not hard to find examples where the compressed bitrate can be several standard deviations above the average bitrate on the training set - see e.g. the last example here: https://github.com/Justin-Tan/high-fidelity-generative-compr...

(Lossy) neural compression methods may also synthesize small portions of an image to avoid compression artefacts associated with standard image codecs, so should definitely not be used in sensitive applications where small details can make a big difference such as security imaging, guarantees or none.

tonic_section··on Differentiable Dithering
Unfortunately you wouldn't have any guarantees on the output of any particular image though, just some reassurances about the expected behaviour over the training set.
tonic_section··on Show HN: Neural Image Compression Demo
Yeah, high frequency detail such as facial features for faraway figures or text tend to get washed out after compression - this is probably due to a couple reasons: 1) The training dataset contains relatively few pictures including text, 2) high-frequency detail is too expensive to encode and the model learns to forgo encoding this in favor of more 'important' features such as shapes, colors, etc.
tonic_section··on Show HN: Neural Image Compression Demo
I think S3 permits up to 20k requests before they start billing IIRC.
tonic_section··on Show HN: Neural Image Compression Demo
Hey, thanks for bringing the brightness issue to my attention - turns out I wasn't normalizing the output correctly - I just pushed a fix and the output images don't have the brightness change now.

As for the random spots, that's an artifact of the entropy coding algorithm. In principle this is lossless but there is some distortion because I'm using a custom vectorized version of an rANS encoder and it's hard to encode overflow values in a vectorized fashion, I'm working on this though. If you can live with really slow decoding times (2-3mins) then you can disable vectorization to eliminate these small imperfections entirely.

As for the comparison to the official model, that's mainly because of compute constraints v. Google (this is just my weekend project). My model uses a smaller architecture and was trained for only 4e5 steps versus the 2e6 steps they reported in the paper - even then it took 4+ days on AWS! The model is also trained on the Openimages dataset, which is presumably much smaller and more noisy than the massive internal dataset Google used.

tonic_section··on Show HN: Neural Image Compression Demo
I eventually shifted the models to S3, but thanks for the offer.
tonic_section··on Show HN: Neural Image Compression Demo
Sorry, looks like both GDrive and Zenodo have exceeded the temporary download quotas, so the model checkpoints aren't available currently... If anyone has any solutions on how to publicly host model weights (~2 GB) please let me know!
tonic_section··on Show HN: Neural Image Compression Demo
I pushed a workaround and provided extra instructions in the demo, so anyone experiencing errors should try that.
tonic_section··on Show HN: Neural Image Compression Demo
GDrive doesn't download the model checkpoints correctly sometimes, leading to the following error:

``` # Setup model

I get an error in the function call 'prepare_model'

UnpicklingError: invalid load key, '<'. ```

Try rerunning the download cell if you experience this - the models downloaded should be around 1.5-2GB, so if the checkpoints are 100kB in size, the download's gone wrong.

tonic_section··on Show HN: Neural Image Compression Demo
Yes, the model is not lossless as this would require learning the PDF in the original input space.

However, the model does learn a conditional probability distribution over a lower-dimensional representation of the original image - this is unavoidable as entropy coding requires a distribution over discrete symbols. The GAN is almost auxiliary and not a central component of the model - in fact, you can get very good results without the GAN, but does seem to result in visually superior reconstructions.

tonic_section··on Show HN: Neural Image Compression Demo
As u/londons_explore mentioned, in theory you can train a model for lossless reconstruction - there are several papers about this, e.g. [1] is a good recent example. Lossless compressors need to learn a probability distribution over each input pixel, which amounts to maximum likelihood estimation in the original image space.

The model in the demo is a lossy compression method because it first projects the input to a lower dimensional space and performs quantization of this representation to integer values so the result can be ultimately entropy coded. It uses the mean-scale hyperprior model introduced in [1] to estimate the necessary probability distributions in the lower-dimensional space for entropy coding.

[1]: https://arxiv.org/abs/1811.12817 [2]: https://arxiv.org/abs/1802.01436

tonic_section··on Show HN: Neural Image Compression Demo
What was the error? I tried to make the demo notebook as robust as possible - you should be able to execute all cells in sequence once then execute cells out of sequence etc. without trouble, but it's hard to legislate for errors in Jupyter-like notebooks sometimes.
tonic_section··on Show HN: Neural Image Compression Demo
The model was trained on a fairly image (~1e6) dataset of diverse high-resolution natural images (the Openimages dataset) - so there was no particular training domain, and generalizes to images of arbitrary size/resolution/content well. There is a larger set of samples generated using the medium bitrate model which can be viewed in this Google Drive: https://drive.google.com/drive/folders/1lH1pTmekC1jL-gPi1fhE...

One interesting failure model is that images dominated by high-frequency detail require a relatively large bitrate to store - see e.g. the last example in the Github README with the weird brickwork. Even though the model was trained to produce compressed representations with a soft constraint on the maximum bitrate, the filesize of the representation for this particular image is something like 60% above the nominal maximum.

tonic_section··on Show HN: Neural Image Compression Demo
During training, you can set a target bitrate by heavily penalizing examples which exceed the target rate in the rate-distortion objective - so the model should learn to produce compressed representations at or below this bitrate. However, this constraint is only enforced on aggregate throughout the entire dataset - like many ML systems, there is no guarantee of behaviour for individual examples, either within or outside the training set. Despite this, the model appears to respect the target rate well, even on out-of-sample images.

One shortcoming is that this current model is non-adaptive - which means that the target rate is fixed. So to achieve different target compression rates you would have to train multiple models in different rate regimes. In the Colab demo there is the option to select between 3 different models trained with a target bits-per-pixel (bpp) rate at 0.14bpp, 0.30bpp, and 0.45bpp, respectively - higher rates correspond to more higher-fidelity reconstructions, at the expense of a lower compression ratio. The default is the `HiFIC-med` model (and this is what the all samples in the README were generated with), but the model trained at the highest bitrate should have less obvious imperfections.

There's also an aspect to the distortion that can be attributed to the entropy coding process rather than the model itself - currently the system clips values outside a certain probability range, resulting in artificial distortion - a fix is in the pipeline though.

tonic_section··on Show HN: Neural Image Compression Demo
Hi everyone, I've been working on an implementation of a model for learnable image compression together with general support for neural image compression in PyTorch. You can try it out directly and compress your own images in Google Colab [1] or checkout the source on Github [2].

This project is based on the paper "High-Fidelity Image Compression" by Mentzer et. al. [3] - this was one of the most interesting papers I've read this year! The model is capable of compressing images of arbitrary size and resolution to bitrates competitive with state-of-the-art compression methods while maintaining a very high perceptual quality. At a high-level, the model jointly trains an autoencoding architecture together with a GAN-like component to encourage faithful reconstructions, combined with a hierarchical probability model to perform the entropy coding.

What's interesting is that the model avoids compression artifacts associated with standard image codecs by subsampling high-frequency detail in the image while preserving the global features of the image very well - for example, the model learns to sacrifice faithful reconstruction of e.g. faces and writing and use these 'bits' in other places to keep the overall bitrate low.

The overall model is around 700MB - so transmitting the model wouldn't be particularly feasible, and the idea is that both the sender and receiver have access to the model, and can transmit the compressed messages between themselves.

If you have any questions or notice something weird I'd be more than happy to address them.

---

[1] Colab Demo: https://colab.research.google.com/github/Justin-Tan/high-fid...

[2]: Github: https://github.com/Justin-Tan/high-fidelity-generative-compr...

[3]: Original paper: https://hific.github.io/

[4]: Sample reconstructions: https://github.com/Justin-Tan/high-fidelity-generative-compr...

tonic_section··on PyTorch for Scientific Computing: Quantum Mechanics Example Part 2
CuPy shares a lot of the Numpy API. I've found it pretty interchangable in most applications.
tonic_section··on Listening for illegal logging chainsaws using TensorFlow
The problem is that neural networks trained using maximum LL do not return calibrated probabilities, using e.g. the softmax output as 'confidence' of a model tends to result in overconfident predictions, take a look at adversarial attacks on neural networks for an extreme example: https://blog.openai.com/adversarial-example-research/
tonic_section··on Listening for illegal logging chainsaws using TensorFlow
How do you quantify the confidence of your model? Do you use a Bayesian model or just the log-likelihood? Because the latter can act strangely in some cases.