Video codec in 100 lines of Rust
blog.tempus-ex.com
blog.tempus-ex.com
QOI https://qoiformat.org/ is a good example of a practically useful simple format.
It's still quite a leap to get to the best new codecs, suddenly you are in a world of head hurty math.
It's also worth noting that it is easy to beat the old formats all-round with off-the-shelf parts, both JPEG and PNG can be bested by changing the outer level of compression for something that was invented after the formats were made. For instance using LZMA or zstd as the final stage. Quite often that's enough to put them on a par with more radically different newer formats.
The heart of JPEG is in The DCT and its energy compaction properties.
One crazy thing about the DCT is that it doesn't just let you make trade-offs for high/low frequency features. It also lets you make tradeoffs for horizontal and vertical features. If you customize your quantization matrix to your specific application, you can potentially achieve compression ratios far exceeding anything available today - even if you leave in the crusty old RLE+Huffman coding.
If you want to get up to your elbows in this sort of thing, there is an entire book on it by Rao & Yip that is about as comprehensive as it gets - https://www.abebooks.com/products/isbn/9780125802031/3119886...
If you do that with a video codec, you lose the ability to seek within the stream, which makes it useless for streaming video. For images, the amount of memory required to decompress may become excessive.
It's similar to why zip (deflate) is still widely used, but is far from optimal in compression efficiency; everything that does better (in some cases much better) is going to be slower, bigger, or both. See https://en.wikipedia.org/wiki/PAQ for an example of extreme lossless compression.
Files themselves are not compressed on btrfs, it's the blocks that get compressed.
ZSTD also has some fun "dictionary" operations so that even if you're chunking your data you can still take advantage of cross-chunk redundancy by "training" across all your chunks before the compression stage.
according to https://cloudinary.com/blog/contemplating-codec-comparisons#... google leaned hard on AVIF ability to produce small garbage in its comparison against jpeg xl
It's interesting to see how well such a simple technique performs. I wonder what would happen if you added trivial temporal compression by simply subtracting the color values of the previous frame from the next and encoding the residual. How would that perform?
Quite a few codecs in the "intra-frame only" section of this Wikipedia list, and that section is within the "Video compression formats" section:
https://en.wikipedia.org/wiki/List_of_codecs#Intra-frame-onl...
I was certainly expecting some motion coding.
Some video formats only go I. Then there's not a lot different between images and video, as far as editing goes. Decoding for end user transportation has a lot more going on, but one has to start somewhere.
Anyway - I think that this kind of work is a great starter and gets more people interested in this.
Instead of working with a delta, conditionally using previous frame as prediction source could work (e.g. if pixel A was closer to previous frame's A than to current frame's B, predict from previous frame's X). Or you could signal prediction source explicitly per block or with RLE. Ideally you'd do motion compensation, but doing that precisely enough for a lossless compressor is more than 100 lines.
Personally, I use ProRes 422 for recording, and DNxHD/DNxHR for proxies (and that's only because DaVinci Resolve's free edition can't create ProRes Proxies).
Both of these codecs are all-intra formats in mpeg containers.
see: mjpeg
Not for H.264; looks like the last patent expires in 2028:
https://scratchpad.fandom.com/wiki/MPEG_patent_lists#H.264_p...
On the other hand, the last patent on MPEG-4 ASP (Xvid/DivX/etc.) which preceded H.264 apparently just expired earlier this month:
https://meta.wikimedia.org/wiki/Have_the_patents_for_MPEG-4_...
...and IANAL but that means the patents for H.263 and everything older should've already expired too.
It's harder to evaluate the blob of patents from 2004-2005.
There’s rav1e for AV1 encoding.
In contrast, you can build a very simple encoder using very few of the format features and still have it be usable/useful (albeit with poor quality/compression ratio).
Weirdly, chatgpt can be remarkably good at translating code between programming languages.
I suspect within a year or two it'll be pretty easy to translate a lot of C libraries to native rust code (or whatever) using modern AIs.