https://storage.googleapis.com/demos.webmproject.org/webp/cm...
Most comparison links posted here are to older (almost a year old) versions that don't reflect the current state of encoding. Both JPEG XL and AVIF have improved tremendously.
https://storage.googleapis.com/demos.webmproject.org/webp/cm...
Most comparison links posted here are to older (almost a year old) versions that don't reflect the current state of encoding. Both JPEG XL and AVIF have improved tremendously.
1) Progressive decoding. Like the original .jpg, .jxl can give you a low-quality image when a fraction of the file is loaded, then a decent-quality image, then the final image. This can give JPEG XL the edge in perceived load speed even when the full .avif is smaller than the full .jxl. (Old demo from a JXL contributor at https://www.youtube.com/watch?v=UphN1_7nP8U )
2) Fast conversion: JPEG XL encoding/decoding is fast without dedicated hardware. Facebook found encode/decode speed and progressive decoding to be points in favor of JPEG XL for their use: https://bugzilla.mozilla.org/show_bug.cgi?id=1539075#c18
3) .jpg repacking: JPEG XL can pack a JPEG1 about 20% smaller without any additional loss; the original .jpg file can be recovered bit-for-bit.
4) Lossless mode. JXL's lossless mode is the successor to FLIF/FUIF, is really good, and also has progressive decoding. AVIF has a lossless mode too, but JPEG XL seems ahead here.
(I know the parent comment is from a JXL contributor, I'm saying this for other folks.)
I think those will give JPEG XL a niche on the Web. Meanwhile I suspect e.g. Android phone cameras will save .avif someday, like iPhones save .heic now. Phones want the encode hardware anyway for video, and you can crunch a zillion megapixels down to a smaller file with AVIF before attention-grabbing artifacts crop up--at low bitrates AVIF seems good at preserving sharp lines and mostly blurring low-contrast details (compare Tiny images).
Finally, worth noting the codecs are different due to a bunch of rational choices by their devs. AVIF is the format for AV1 video keyframes. Progressive decoding doesn't help there, and doesn't jibe well with spatial prediction, which helps AV1 and other video codecs preserve sharp lines. And video codecs need hardware support to thrive anyway, so optimizing for fast software encoding probably wasn't an early priority. Otherwise the new formats have a lot of overlap in fundamentals--variable size and shape DCTs, better entropy encoding, chroma-from-luma, anti-ringing postfilters, etc.
Glad to see support for both getting more widespread.
Note this means that animated images on the web (like GIF) are significantly smaller with AVIF than JPEG-XL which has no inter prediction.
JPEG XL does have some weak forms of inter prediction though (but they were designed mostly for still image purposes). One of them is patches: you can take any rectangle from a previously 'saved' frame (there are four 'slots' for that) and blit it at some arbitrary position in the current frame, using some blend mode of choice (just replace, add, alpha blend over, multiply, alpha blend under, etc). This is obviously not as powerful as full motion vectors etc, but it does bring some possibilities for something like a simple moving sprite animation. This coding tool is currently only used in the encoder for still images, namely to extract repeating text-like elements in an image (individual letters, icons etc) and store them in a separate invisible frame, encoded with non-DCT methods (which are more effective for that kind of thing) and then patch-add them to the VarDCT image. The current jxl encoder is not even trying to be good at animation because this is not quite its purpose (it can do it, but 'reluctantly').
Anyway, I think that animation is in any case best done with video codecs (this is what video codecs are made for), and I wish browsers would just start accepting in an <img> tag all the video codecs they accept in a <video> tag (just played looping, muted, autoplay), so we can for once and for all get rid of GIF.
Any format that doesn't have this is doomed to fail as a GIF replacement.
Also a plus for saving phone snaps, since the camera often saves a short video these days anyway.
At 'Large' and 'Big' settings of this image -- which are still in much less than 1 bpp bitrates, i.e., below internet image quality -- you can still observe significant differences in the clouds even if balloons are relatively well rendered.
JPEG XL is the first codec to have a practical encoder that can be configured by saying "I want the worst visual difference to be X units of just-noticeable-difference". All other encoders are basically configured by saying "I want to use this scaling factor for the quantization tables, and let's hope that the result will look OK".
crf in x264/x265 is smarter than that, but it's still a closed-form solution. That's probably easier to work with than optimizing for constant SSIM or whatever, it always takes one pass and those objective metrics are not actually very good.
It is a bit like looking at bitrate for Video quality without looking at video resolution.
[1] https://en.wikipedia.org/wiki/JPEG_XL#Standardization_status
Part 3 will describe conformance testing (how to verify that an alternative decoder implementation is in fact correct), and part 4 will just be just a snapshot of the reference software that gets archived by ISO, but for all practical purposes you should just get the most recent git version. Parts 3 and 4 are not at all needed to start using JPEG XL.
Bitrates vary from 0.26 bpp (Nestor/AVIF) to 4+ bpp (205/AVIF) at the finest setting. Nestor at lowest setting is just 0.05 bpp, somewhat unusual for an internet image. A full HD image at 0.05 bpp transfers over average mobile speed in 5 ms and is 12 kB in size. I rather wait for a full 100 ms and get a proper 1 bpp image.
I just compared the original tiger image to the lossless version for JPEG XL but there's some small changes that it makes.