The Unreasonable Effectiveness of JPEG: A Signal Processing Approach [video]
youtube.com
youtube.com
I feel the focus on the beauty and amazement is somewhat misleading. It's like a magic trick, they're trying to fool you into thinking this is some superman feat.
I want the story of the JPEG told from the "let's just throw some information away because it's good enough to get the current job done even if it looks crap" perspective.
Like reading about how your favourite singer or comedian grew up doing impersonations of their favourite singer or comedian, I feel it gives a truer glimpse into how these things work in reality and a better understanding for people who want to learn.
Engineering is all about comprise and sacrifice to get a product working within the constraints (price, weight, volume, time to market, project risk, regulatory compliance, reliability, etc), so these stories would exist for everything.
Except Juicero, that thing was engineered to death: the only thing they (well...the VCs) skimped on was the market research and common sense.
The thing built like that would either be crap and nobody uses it so no cool story, or it will be so damn genius with middle out that gets a Weissman score of 9.8 or some such.
https://en.wikipedia.org/wiki/JPEG#Background
> The original JPEG specification published in 1992 implements processes from various earlier research papers and patents cited by the CCITT (now ITU-T) and Joint Photographic Experts Group.[1] The main basis for JPEG's lossy compression algorithm is the discrete cosine transform (DCT),[1][8] which was first proposed by Nasir Ahmed as an image compression technique in 1972.[9][8] Ahmed developed a practical DCT algorithm with T. Natarajan of Kansas State University and K. R. Rao of the University of Texas in 1973.[9] Their 1974 paper[15] is cited in the JPEG specification, along with several later research papers that did further work on DCT, including a 1977 paper by Wen-Hsiung Chen, C.H. Smith and S.C. Fralick that described a fast DCT algorithm,[1][16] as well as a 1978 paper by N.J. Narasinha and S.C. Fralick, and a 1984 paper by B.G. Lee.[1] The specification also cites a 1984 paper by Wen-Hsiung Chen and W.K. Pratt as an influence on its quantization algorithm,[1][17] and David A. Huffman's 1952 paper for its Huffman coding algorithm.[1]
The JPEG2000 standards major move forward was using wavelets instead of cosines, building on a different strand of research also going back decades...
1) Convert image to YUV color space so that brightness (Y) component can be compressed separately from color (UV) components to which humans are less sensitive.
2) Transform each component (YUV) to frequency domain, then throw away image fine detail, i.e. high frequency components, according to desired level of compression. This obviously is the key idea.
3) Encode remaining frequency components to as small a size a possible
Once the general idea for this approach was conceived, it seems a rather minor step to try DCT vs the more obvious FFT for the frequency encoding. DCT gives slightly better compression. The approach isn't dependent on Huffman encoding for step 3) either - any encoding scheme would work, so experimentation would also have worked there to see what gives best compression on some set of test images.
I've got implementations on top of libjpegturbo that can encode with latency low enough to support interactive frame rates at modern resolutions. We are talking 1920x1080@4:4:4 in under 10 milliseconds.
Then for some weird reason it gained traction as a "good quality video codec" (probably because there's no motion vectors, hence the associated artifacts simply can't happen with it).
If you want throughput, hardware-based h.264 gonna be faster than JPEG while delivering better quality. It’s possible to configure the encoder so it only emits keyframes (I-frames), which makes the video stream a sequence of independent compressed frames. However don’t forget about parameter sets, i.e. SPS/PPS blobs – the decoder gonna need them for each frame.
If you want latency, consider BC1 https://docs.microsoft.com/en-us/windows/win32/direct3d10/d3... or BC7 https://docs.microsoft.com/en-us/windows/win32/direct3d11/bc... They are lossy codecs usually used for textures in 3D scenes, the decoders are implemented in hardware in all modern PC GPUs. The algorithms are way simpler than JPEG, also easily parallelizable. Compression ratio is much worse though, for RGB24 images it’s 33% for BC7 and 17% for BC1, JPEG compression is way more efficient.