The dangers behind image resizing (2021)
zuru.tech
zuru.tech
1. The form of interpolation (this article).
2. The colorspace used for doing the arithmetic for interpolation. You most likely want a linear colorspace here.
3. Clipping. Resizing is typically done in two phases, once resizing in x then in y direction, not necessarily in this order. If the kernel used has values outside of the range [0, 1] (like Lanczos) and for intermediate results you only capture the range [0,1], then you might get clipping in the intermediate image, which can cause artifacts.
4. Quantization and dithering.
5. If you have an alpha channel, using pre-multiplied alpha for interpolation arithmetic.
I'm not trying to be exhaustive here. ImageWorsener's page has a nice reading list[1].
Here is imageworsener's article about this[1]
Here is the most popular article about this problem [1].
Warning: once you start noticing incorrect color blending done in sRGB space, then you will see it everywhere.
It is unquestionably superior though.
Upscaling is much more difficult.
>> The definition of scaling function is mathematical and should never be a function of the library being used.
I could just as easily say "hey, why is you NN affected by image artifacts, isn't it supposed to be robust?"
Also, gamut clipping and interpolation[0]. That's a real rabbithole.
[0] https://www.cis.rit.edu/people/faculty/montag/PDFs/057.PDF (Downloads a PDF)
One pet peeve of mine is algorithms for making thumbnails, most of the algorithms from the image processing book don't really apply as they are usually trying to interpolate between points based on a small neighborhood whereas if you are downscaling by a large factor (say 10) the obvious thing to do is sample the pixels in the input image that intersect with the pixel in the output image (100 in that case.)
That box averaging is a pretty expensive convolution so most libraries usually downscale images by powers of 2 and then interpolate from the closest such image which I think is not quite perfect and I think you could do better.
Practically people think that box averaging is too expensive (pretty much it is like that Gaussian blur but computed on fewer output points.)
[1]: https://www.intel.com/content/dam/develop/external/us/en/doc...
If you shrink 10x in one direction, then the other, then you first turn 100 pixels into 10, before turning 10 pixels into 1. You actually do more work for a non-smoothed shrink, sampling 110 pixels total.
To benefit from doing the dimensions separately, the width of your sample has to be bigger than the shrink factor. The best case is a blur where you're not shrinking at all, and that's where 20:1 actually happens.
If you sampled 10 pixels wide, then shrunk by a factor of 3, you'd have 100 samples per output if you do both dimensions at the same time, and 40 samples per output if you do one dimension at a time.
Two dimensions at the same time need width^2 samples
Two dimensions, one after the other, need width*(shrink_factor + 1) samples
Of course if you're not interpolating but downscaling the image (which isn't really an interpolation, the value at a particular position in the image does not remain the same) then you do want a linear colourspace to avoid brightening / darkening details, but you need a perceptual colourspace to minimize ringing etc. It's an interesting puzzle.
If one is concerned about this, one could intentionally vary the resampling or deliberately add different blurring filters during training to make the model robust to these variations
I’ve seen it cause trouble in every model architecture i’ve tried.
it’s not a huge instability, but you can absolutely see performance changes.
I say that choice of resampling algorithm is what determines whether a model can learn the rule “zebras can be recognized by their uniform-width stripes” or not; as a bad resample will result in non-uniform-width stripes (or, at sufficiently small scales, loss of stripes!)
Unless it's actually making a visible change that spoils whatever it is the model supposed to be looking forBut zebras don't have uniform-width stripes. https://www.animalfactsencyclopedia.com/Zebra-facts.html
Other supposedly better CUDA/ML filters give me strange results.
I really wish there are some better general-purpose imaging libraries that steadily implement/copy these useful filters, so that more people can use them out of the box.
Most of languages I've involved are surprisingly lacking in this regard despite their huge potential use cases.
Like, in case of Python, Pillow is fine but it has nothing fancy. You can't even fine-tune parameters of bicubic, let alone billions of new algorithms from video communities.
OpenCV or ML tools like to re-invent the wheels themselves, but often only the most basic ones (and badly as noted in this article).
A big sticking point is variable resolution, which it technically supports but doesn't really like without some workarounds.
But yeah I agree, its kinda tragic that the ML community is stuck with the simpler stuff.
I found https://dl.acm.org/doi/10.1145/2766891 but I don't like the comparisons. Any designer will tell you, after down-scaling you do a minimal sharpening pass. The "perceptual downscaling" looks slightly over-sharpened to me.
I'd love to compare something I sharpened in photoshop with these results.
clip = core.imwri.Read(img)
clip = muf.ssim_downscale(clip, x, y)
clip = core.imwri.Write(clip, imgoutput)
clip.set_output()
> Any designer will tell you, after down-scaling you do a minimal sharpening pass
This is probably wisdom from bicubic scaling, but you usually dont need further sharpening if you use a "sharp" filter like Mitchell.
Anyway I havent run butteraugli or ssim metrics vs other scalers, I just subjectively observed that ssim_downscale was preserving some edges in video frames that Spline36, Mitchell, and Bicubic were not preserving.
Horseshit. Image resizing or any other kind of resampling is essentially always about filling in missing information. The is no mathematical model that will tell you for certain what the missing information is.
That's why a vector image rendered at 128x128 can look better/sharper than one rendered at 256x256 and scaled down.
In your example the lower res image would be using most of its bandwidth while the higher res image would be using almost none of its bandwidth.
Images are 2D discrete signals. Everything you know about 1D DSP applies to them.
Take the naive case where you downscale a line of four pixels to two pixels - you can simply discard two of them so you go from `0,1,2,3` to `0,2`. It looks okay.
But what happens if you want to scale four pixels to three? You could simply throw one away but then things will look wobbly and lumpy. So you need to take your four pixels, and fill in a missing value that lands slap bang between 1 and 2. Worse, you actually need to treat 0 and 3 as missing values too because they will be somewhat affected by spreading them into the middle pixel.
So yes, downscaling does have to compute missing values even in your naive linear interpolation!
Like yeah, you can try to get clever and preserve the artistic intent or something with something like seamcarving but then I wouldn't call it downscaling anymore.
This is already wrong, unless the pixels are band-limited to Nyquist/4. Trivial example where this is not true:
1 0 1 0
If such a signal is decimated by 2 you get 1 1
Which is not correct.And besides, a ranty blog post pointing out pitfall can still be useful for someone else coming from the same naïve (in a good/neutral way) place as the author.
An example used in the article: https://en.wikipedia.org/wiki/Lanczos_resampling
An anamorphic lens (optically) "squeezes" the image onto the sensor, and afterwards the digital image has to be "desqueezed" (i.e. upscaled in one axis) to give you the "final" image. Which in turn is downscaled to be viewed on either a monitor or a printout.
But the resulting images I've seen until now nevertheless look good. I think that's because in natural images you have not that many pixel-level details. And we mostly see downscaled images on the web or in youtube videos most of the time ...
By that I mean, I know what bilinear/bicubic/lanczos resizing algorithms are, and I know they should at least have acceptable results (compared to NN).
But I don't know famous libraries (especially OpenCV which is a computer vision library!) could have such poor results.
Also a side note, IIRC bilinear and bicubic have constants in the equation. So technically when you're comparing different implementations you need to make sure this input (parameters) is the same. But this shouldn't excuse the extreme poor results in some.
Not so fast: https://entropymine.com/imageworsener/bicubic/
This isn't necessarily a criticism of OpenCV, often the OpenCV implementation is, of necessity, quite general, and a specific use-case can engage optimizations not available in the general case
> Since we noticed that the most correct behavior is given by the Pillow resize and we are interested in deploying our applications in C++, it could be useful to use it in C++. The Pillow image processing algorithms are almost all written in C, but they cannot be directly used because they are designed to be part of the Python wrapper. We, therefore, released a porting of the resize method in a new standalone library that works on cv::Mat so it would be compatible with all OpenCV algorithms. You can find the library here: pillow-resize.
Ooops. Just thought about generative systems. Nevermind.
You can use this to your advantage by purposely introducing them into the lowres inputs so they will be removed.
The comparison of resizing algorithms is not something new, importance of adequate input data is obvious, difference in image processing algorithms availability is also understandable. Clickbaity.
And he was hit by a truck.
So it's true about the danger of image resizing.
I wonder why it's not adopted by any of these frameworks?
IS ML sort of like a universal hologram in that respect?
Zimg is a gold standard to me, but yeah, you can get better output depending on the nature of your content and hardware. I think ESRGAN is state-of-the-art above 2x scales, with the right community model from upscale.wiki, but it is slow and artifacty. And pixel art, for instance, may look better upscaled with xBRZ.
E.g. looks totally different on original vs 224x224 pictures
https://embracethered.com/blog/posts/2020/husky-ai-image-res...
https://pytorch.org/docs/1.9.0/generated/torch.nn.functional...