Decoding AVIF: Deep dive with cats and imgproxy
evilmartians.com
evilmartians.com
AVIF/HEIF has hundreds of features that decoders are theoretically supposed to implement. Many of ISOBMFF "boxes" also have multiple versions (typically 16/32/64-bit versions or "oops we forgot to add a field" version). Does every decoder really need to support all of them? It's such a waste of effort, and file bloat. Browsers only care about getting pixels on screen. They don't have UIs nor APIs to browse through photo bursts, bracketed exposures, or an infrared channel (or is every AVIF viewer supposed to implement these now?)
And some features are even dangerous: lossless cropping. Great for your camera roll app, but for publishing on the Web it's a huge footgun. You could think you've cropped private info out of the picture you've shared, but randos on the internet can uncrop your HEIF pictures.
I feel like I would need a full hour long video on each paragraph of this post to really understand it.
Keyframes were just jpegs. Then for intraframes, you first found the motion vectors for each 8x8 block, then generated the predicted intraframe from the keyframe and motion vector. Then you subtracted the prediction off the keyframe. The resulting "prediction error image" was then simply jpeg encoded as the intraframe, and appended to the output after the motion vectors.
Decoding was reverse, reconstruct the predicted intraframe from the previous keyframe and motion vectors, and add back the prediction error image.
Might be glossing over something as it was over a decade ago but should be about the gist of it.
We played with different algorithms for finding motion vectors and such, including accelerating it with GPGPU.
Really fun project, and assuming you use an existing jpeg library, not at all big or difficult.
In case it wasn't obvious, the error image can of course have negative values which jpeg can't handle. So you add a bias of +128 and clamp the biased error to [0, 255].
During decoding you simply subtract the bias when adding back the error image.
You can find my deobfuscation and detailed explanation of his program here: http://eastfarthing.com/blog/2020-09-14-decoder/
https://en.wikipedia.org/wiki/Chroma_subsampling nicely explains this clever hack's continuing relevance.
TLDR: Chroma subsampling (as done in the 4:2:2 Y'CbCr color space) is used to improve video encoding efficiency by taking advantage of the "human visual system's lower acuity for color differences than for luminance".
Which led to a lot of simple and professional HW/SW being designed for packed 4:2:2, so codecs support 4:2:2 to fit into professional pipelines.
BTW this cat should now and forever be called Lena.
the first is always 4, don’t ask me why"
Wow. Someone going to this much detail on explaining how a video/image codec works, and cannot bother learning what the numbers of chroma subsampling mean?The first number represents the luminance.[0] Even if they know the first number represents luminance, the "don't ask me why" is just horrible on its own. The detail in the image is preserved through the luminance channel. The subsampling in the chroma is much less perceptable to humans, but more more noticeable in the luminance. Therefore, some very smart people learned to cheat the data saved for chroma, but not the luminance. "don't ask me why" in detailed write ups is just bad in so many ways.
"Now, let's break it down the differences between 4:4:4; 4:2:2 and 4:2:0:
The number of pixels that share color is determined by what type of chroma subsampling it is. Each sample is defined by a block of 8 pixels. The first number refers to the size of the sample and its pattern, which is typically 4 pixels wide. The second number refers to how many pixels in the top row will receive color or chroma sampling. The third number shows how many pixels on the bottom row will receive chroma samples"[0]
The block sizes and sub-sampling methods are also why there are warnings issued when trying to scale an image when the dimensions are not divisible by the block sizes. If you try to scale to an odd number, then the sampling within the blocks is broken. If you scale to a number not divisible evenly by the largest block sizes requested, then you also get issues.
[0] https://blog.westpennwire.com/what-is-chroma-subsampling
I assume it is because with 4 you have 3 different subsampling ratios (if you want to keep factors of two, which you typically want to keep algorithms simple)
Okay, the real reason, as far as the bundles of paper I have* is accurate, is that digital chroma subsampling was first invented for MUSE, a Japanese analogue HD video standard (with pre-broadcast digital components). They chose four for horizontal because it's relatively easy to manipulate using their digital systems at the time and two for vertical so that it's easy to handle interlacing stuff. Unfortunately, I'm not Sony or NHK so I can't say for certain why not eight or any other powers of two. Also, Americans (aka the SMPTE) set the 1,080 lines (the Japanese standard is 1,025), the 16:9 compromise (between the European and Japanese 15:9 and cinema 21:9) and the "limited RGB" dilemma that is experienced in digital video systems (that's literally from the days of NTSC signalling!). Both the Japanese NHK/Sony MUSE system and the British IBA (adopted as European) D-MAC system uses the full-range 8-bit system that is used for JPEG (pre-broadcast to analogue, of course).
Analogous to this, the reason why CD audio is 44,100 Hz is because that's the commonality between NTSC (System M, 525-line 480-visible 60-Hz) and 625-line (576-visible 50-Hz) systems. Digital audio was literally stored on U-matic systems at the time, and it was initially only 14-bit PCM rather than the 16-bit PCM of CDs.
* or rather, my employer's mini-library.
I "surely" look forward to your Show HN write up on your new compression algorithm. We've been iteratively getting better at compression for some time now. It seems like everytime it looks like we've wrung every bit out of DCT, someone comes up with some a little more clever. Wavelets looked promising, but never took off.
>why does 480p upsampled ever look better than 1080p at the same bitrate
That's a very vague question. Are you stating that you think 480p upsampled to 1080p at 1.5Mbps looks better than a source at 1080p at 1.5Mbps? I have a hard time believing this to be true.
To understand why the chroma is sub-sampled and not the luminance has to do with how the cones/rods in the eyes work. There's a lot of things you can get away with (or trick if you will) the brain in what it is seeing. Is it better to lose half the height or half the width? Is it better loose more red than green or blue?
"The commonly used leading digit of 4 is a historical reference to a sample rate roughly four times the NTSC or PAL color subcarrier frequency; the notation originated when subcarrier-locked sampling was under discussion for component video. Upon the adoption of component video sampling at 13.5 MHz, the first digit came to specify luma sample rate relative to 3 3⁄8 MHz. HDTV was once supposed to be described as 22:11:11! Since then, the leading digit has – thank-fully – come to be relative to the sample rate in use. Until recently, the initial digit was always 4, since all chroma ratios have been powers of two – 4, 2, or 1. However, 3:1:1 subsampling has been commercialized in an HDTV production system (Sony’s HDCAM), so 3 may now appear as the leading digit. By convention, a leading digit of 2 is never used."
And here is lots of detailed history: https://tech.ebu.ch/docs/techreview/trev_304-rec601_wood.pdf , including a lot of debate in the late 70's about "three-times sub-carrier (3fsc) versus four-times sub-carrier (4fsc) sampling." The victory for team "4" is, I think, why that's the leading digit, even though they ended up compromising on not-quite-4 in the end.