Bilinear down/upsampling, aligning pixel grids, and that infamous GPU half pixel (2021)
bartwronski.com
bartwronski.com
Very cool effect, and one of my absolute favorite games.
[1]: https://en.wikipedia.org/wiki/Mitchell%E2%80%93Netravali_fil...
The ideal is trilinear sampling instead of bilinear. Easy to do with GPU APIs, in D3D11 it’s ID3D11DeviceContext.GenerateMips API call to generate full set of mip levels of the input texture, then generate each output pixel with trilinear sampler. When doing non-uniform downsampling, use anisotropic sampler instead of trilinear.
Not sure high level image processing libraries like the mentioned numpy, PIL or TensorFlow are doing anything like that, though.
Trilinear interpolates across three dimensions, such as 3D textures or mip chains. It is not a method for downsampling, but a method for filtering that interpolates two bilinear results, such as two bilinear filters of mip levels that were generated with "some" downsampling filter (which can be anything from box to Lanczos).
Anisotropic is a hybrid between trilinear across a smaller interpolation axis under perspective projection of a 3D asset and multiple taps along the longer axis. (More expensive)
I meant trilinear interpolation across the mip chain.
> generated with "some" downsampling filter (which can be anything from box to Lanczos)
In practice, whichever method is implemented in user-mode half of GPU drivers is pretty good.
> It is not a method for downsampling
No, but it can be applied for downsampling as well.
> under perspective projection of a 3D asset
Texture samplers don’t know or care about projections. They only take 2D texture coordinates, and screen-space derivatives of these. This is precisely what enables to use texture samplers to downsample images.
The only caveat, if you do that by dispatching a compute shader as opposed to rendering a full-screen triangle, you’ll have to supply screen-space derivatives manually in the arguments of Texture2D.SampleGrad method. When doing non-uniform downsampling without perspective projections, these ddx/ddy numbers are the same for all output pixels, and are trivial to compute on CPU before dispatching the shader.
> More expensive
On modern computers, the performance overhead of anisotropic sampling compared to trilinear is just barely measurable.
If you want to downsample 3x or fractional, then interpolating between two MIP levels is gonna be worse quality than directly sampling the original image.
Perspective (the use case for anisotropic filtering) isn't discussed in the article, but even then, the best quality will come from something like an EWA filter, not from anisotropic filtering which is designed for speed, not quality.
Have you never downscaled and upscaled images in a non-3D-rendering context?
Indeed, and I found that leveraging hardware texture samplers is the best approach even for command-like tools which don’t render anything.
Simple CPU running C++ is just too slow for large images.
Apart from a few easy cases like 2x2 downsampling discussed in the article, SIMD optimized CPU implementations are very complicated for non-integer or non-uniform scaling factors, as they often require dynamic dispatch to run on older computers without AVX2. And despite the SIMD, still couple orders of magnitude slower compared to GPU hardware while only delivering barely observable quality win.
It has worse quality than something like a Lanczos filter, and it requires computing image pyramids first, i.e., it is also slower for the very common use case of rescaling images just once. And that article isn't really about projected/distorted textures, where trilinear filtering actually makes sense.