How are images compressed? An explanation of JPEG [video]
youtube.com
youtube.com
This video was extremely helpful to understand the "why" of all the things the spec was trying to explain. It made a huge difference in us being able to get things working.
We talk a bit about JPEG and actually writing our decoder in Nim here: https://www.youtube.com/watch?v=vYwD7OynFcg
Overall, our concluding opinion is that JPEG has some extremely cool and really smart ideas for how to compress images but the binary file format itself has some very painful things in it (progressive and restart markers as a couple examples).
JPEG XL appears to finally be getting there, with a meaningful improvement. But still nothing like 3x. Also perhaps AVIF, but current encoders have problems with rough texture on high bitrates.
Some may argue about lossy audio format (as it the nineties called and wanted their 20 GB HDD back) while others may argue about SACD, 24 bit 96 kHz and whatnots but the fact stays: there are engineers out there who came up with the CD audio format in 1980, which is still in use to this day.
I legally and bit-perfectly rip my CDs to FLAC files and it still boggles my mind that it's basically the format from 1980 (FLAC files are lossless and you can re-burn the exact same identical CD, which you can then, if you fancy so, re-rip to the exact same, bit perfect, WAV or FLAC files, rinse & repeat as many times as you want).
Speakers definitely got better. DACs are ubiquitous now. Amps probably got better too. But 16-bit 44.1 kHz stereo lives on since 42 years (40 years commercially). Soon half a century.
"It's like alien technology from the future" indeed.
As far as the xiph.org audio codecs go however, Opus is the real magnum opus (pun obviously intended). SILK (the LPC part, donated by skype) + CELT + DNN (used to detect whether it's speech or music to tune the 2 codecs since libopus v1.3), it's quite complex, and I feel like some of its parts (specifically the SILK encoder, which has the donated implementation and only the high level details in its RFC, since CELT has a plethora of documentation/articles and independent encoder re-implementation in ffmpeg) are only really understood by the original authors (or at least were when they wrote them a decade and a half ago). Reverse engineering the (SILK) encoder code and making a video similar to the one on the OP (or at least an article/blog post) could be a fun activity.
The new kids on the block in the speech encoding/real time communications space (Google Lyra/Microsoft Satin) have fancy AI models, promise decent quality in ultra-low bitrates (3-6kbps), but don't look like they're any easier to run on micro controller.
I think that was the early 2000s. The nineties were the era of the CD-ROM storing huge games that did not fit on your HDD.
edit: is that perhaps how progressive jpegs work and is that why they typically compress better?
Is the chrimonance quantization done for both red and blue chrominance?
I just did something like this! I was messing around and used SVD to write a compressor. Do I dare share my feeble GitHub here lol