Nasir Ahmed's digital-compression breakthrough helped make JPEGs/MPEGs possible
spectrum.ieee.org
spectrum.ieee.org
1. The DCT (II) packs lower frequency coefficients into the top-left corner of the block.
2. Quantization helps to zero out many higher frequency coefficients (toward bottom-right corner). This is where your information loss occurs.
3. Clever zig-zag scanning of the quantized coefficients means that you wind up with long runs of zeroes.
4. Zig-zag scanned blocks are RLE coded. This is the first form of actual compression.
5. RLE coded blocks are sent through huffman or arithmetic coding. This is the final form of actual compression (for intra-frame-only/JPEG considerations). Additional compression occurs in MPEG, et. al. with interframe techniques.
X265/HEVC https://en.m.wikipedia.org/wiki/High_Efficiency_Video_Coding
Also not true for X266/VVC.
https://fgiesen.wordpress.com/2013/11/04/bink-2-2-integer-dc... has a ton of technical information if you want to dive into it.
[edit]
Combined with the information from sibling comments, it seems that the Hadamard transform was something used in standards developed in the '00s but not since.
1. It wouldn't have support for DCTs not in JPEG.
2. It wouldn't use the DCT in its lossy-compression of photographic content if another transform was considered significantly better.
Perhaps one could argue that they didn't want to add extra transforms, but they do use a modified Haar transform for e.g. synthetic content and alpha channels.
What is interesting is that the techniques that are being used in compression, communication and signal processing in general involved orthogonality and Nasir's master and PhD research thesis were in the area of Orthogonal Transform for Digital Signal Processing.
Hadamard Transform can provide orthogonality but unlike DCT and DFT/FFT that are limited to real and complex respectively, Hadamard Transform is very versatile and can be used in real, complex, quaternion and also octonion numbering schemes that probably the latter are more suited for higher dimensions data and signal processing.
Hadamard orthogonal codes has also been used as ECC in reliable space communication in both the Mariner and Voyager missions, for examples [1].
[1] On some applications of Hadamard matrices [PDF]:
https://documents.uow.edu.au/~jennie/WEB/WEB05-10/2005_12.pd...
You can also look at 90's software video codecs developed when DCT was still too expensive for video. They had all kinds of approaches to quantization and entropy coding, and they all were a pixelated mess.
DCT is the key ingredient that enabled compression of photographic content.
The main idea of lossy image compression is throwing away file detail, which means converting to frequency domain and throwing away high frequency coefficients. Conceptually FFT would work fine for this, so use of DCT instead seems more like an optimization rather than a key component.
This matches much better with what happens when you cut out small pieces of a signal, so it gives less noise and thus better energy isolation. Say your little pixel block (let's make it 8x1 for simplicity) is a simple gradient, so it goes 1, 2, 3, 4, 5, 6, 7, 8. What do you think has the least amount of high-frequency content you'd need to deal with; 123456788765432112345678… or 123456781234567812345678…?
(There's also a DCT version that goes more like 1234567876543212345678 etc., but that's a different story)
So still useful for compression, just in more specialized circumstances.
So perhaps it would fair to give due credit to the co-workers as well.
> (subtitle) His digital-compression breakthrough helped make JPEGs and MPEGs possible
Technically, the DCT isn't restricted to only digital compression. The DCT performs a matrix multiplication on a real vector, giving a real vector as output. You can perform a DCT on a finite sequence of analog values if you really wanted to, by performing a specific weighted sum of the values to yield a new sequence of analog values.
My graduate thesis advisor was a coinventor of the DCT [0]. I miss my grad school days - he was a great advisor.
Having files around which, for twenty years, couldn't be compressed losslessly and that now suddenly can is just wild.
And even though I didn't look that much into it JPEG XL is, basically... More DCT!?
Consider a shallow 1D gradient, which is just a ramp. The DFT's interpretation of this as a periodic signal turns it into a sawtooth, which takes lots of high frequency components to reduce ringing and keep the edge sharp enough. The DCT is equivalent to the DFT on the mirrored signal, which instead turns this into a triangle wave, which takes less high frequency components to represent reasonably.
This varies for different types of data; my understanding is that audio codecs tend to prefer the modified DCT (MDCT) instead due to the different characteristics of audio signals.
Is that appropriate? Would another discipline give me a better grounding in not just these techniques, but the mental foundations that made their discovery possible?
Could you recommend some of the books you mentioned?
Equations, differentiation, integration, partial equations, complex numbers, matrices, eigen vectors, correlation, Fourier transform, laplace transform. And probably others.
I find the relationship between compression algos & cognitive science very interesting.
Lossless compression is entirely reversible. Nothing is lost and nothing is discarded, perceived or not, like zip.
Lossless is formats like Flac and zip. Lossless compression basically stores the same data in more efficient (from a file size perspective) states rather than discarding stuff that isn’t perceived.
The clue is in the name of the term: “lossy” means you lose data. “Lossless” means you don’t lose data. So if a zip file was lossy, you’d never be able to decompress it. Whereas you cannot restore data you’ve lost from an MP3.
But this is an example https://archive.is/KShWY#9
I am only aware of JPEG, actually. Can anyone help me? PNG uses deflate and not DCT, TIFF supports sort of everything (JPEG is uncommon but possible) but generally no DCT is used, GIF uses some RLE but also not with a DCT, J2K also does not use a DCT, EXR can use Wavelet as well but no DCT I'm aware of.
JPEG-XL also uses DCT.
As an aside, while PNG uses deflate for entropy coding, the conceptual analog to DCT in the context of PNG would be its row filters. JPEG's entropy coding isn't all that different to PNG's (aside from the arithmetic coding option which isn't widely used).
And yes, I am sorry mixing DCT with entropy coding, I've noticed already during writing my comment, but didn't find a better way, and I see you understood what I meant.