JPEG XL was based (in part) on Google's "pik" codec, which used the Butteraugli metric for bitrate targeting.
Use your eyes instead: https://afontenot.github.io/image-formats-comparison/#wasser...
JPEG XL was based (in part) on Google's "pik" codec, which used the Butteraugli metric for bitrate targeting.
Use your eyes instead: https://afontenot.github.io/image-formats-comparison/#wasser...
Non-blind n=1 human test can be just as flawed. What you like in a particular scenario is not representative of codecs' overall performance.
Testing with humans requires proper setup and a large sample size (which BTW JPEG XL has done!)
The problem is that these codecs are close enough in performance to be below "noise" level of human judgement.
You will not be able to reliably distinguish q=80 vs q=81 of the same codec, even though there is objectively a difference between them. And you can't lower quality to potato level to make your job easier — that changes the nature of the benchmark.
People also just differ in opinion whether blurry without detail is better than detailed but with blocking or ringing. If you ask people to rate images, stddev of their scores will be pretty large, so wide that scores of objective metrics can fit in 1-sigma.
People also tend to pick "nicer" image, rather than the one that is closer to the original. That's a bit different task than codecs aim for.
Codecs can allocate more or less of the file to color channels, so you can get different conclusions based on e.g. amount of bright red in the image.
So testing is hard. Plenty of pitfalls. Showing a smaller file that "looks the same" is easy, but deceptive.
> Showing a smaller file that "looks the same" is easy, but deceptive.
It's much better to show same-sized files and let the viewer assess their quality (this is what the comparison I linked does), but there are deceptive ways to do even this.
See also this great old blog post about doing comparisons for video codecs: https://web.archive.org/web/20141103202912/http://x264dev.mu...
Samsung turns their screens and phone cameras to “vivid” processing by default (they do provide a “natural” toggle, to their credit).
I think over 90% of people don’t realize what they are seeing is not close to reality, and instead ask people with other phones, especially iPhones, why their screen or photo looks so washed out.
There’s many faults to BFDLs, but I love Apple providing “vivid” processing as an option and “natural” as the default. Sometimes the masses just don’t know what’s actually good for them.
For instance, at medium quality Jxl seems to be better at preserving fine details and structure like the mark on the door, the traces on the lower part of the bridge, but avif appears to be better at preserving clarity of complex details, like far-away windows, cars, a tennis racket.
Sharp edges are just one texture in an infinite range of textures, and AVIF looks like it constructs everything out of sharp edges in a way that's really obvious to the eye at all compression levels.
With JPEG XL I can at least tell what's missing, or too artifacted to make out. With AVIF you have no idea what has been completely erased.
In case you actually wanted to compress images specifically for viewing when zoomed in, you should use different codecs, or higher quality, or configure codecs differently (e.g. in classic JPEG make quantization table preserve high frequences more).
But for a benchmark that claims to compare codecs in general, you should only use normal viewing distance. Currently it's controversial whether the norm is still ~100dpi or whether it should be the "Retina"/2x resolution, but definitely it's not some zoomed-in 5dpi.
https://afontenot.github.io/image-formats-comparison/#vid-ga...
It's visible even in qualities higher than "tiny", which IMO is unreasonably small.
Frankly I can't really see much difference between AV1/Tiny and JPEGXL/Large on the [1] middle of the coffee one, everything else around (cup, receipt, spoon) clearly have more detail but not the coffee in the middle
https://afontenot.github.io/image-formats-comparison/#catedr... (glasswork)
https://afontenot.github.io/image-formats-comparison/#end-of... (teeth, fingernails)
https://afontenot.github.io/image-formats-comparison/#steinw... (lights, pianos)
https://afontenot.github.io/image-formats-comparison/#us-ope... (racket, fingernails, jewelry, watch)
these features are non-helpful at normal photography bitrates and only complicate coding at bitrates above 1.5 or so
JPEG XL has similar approaches but its tools have a larger quality operating range
We evaluated these tools for JPEG XL and I rejected them due to them only helping at very lowest bit rates
there are many other ideas on how low quality JPEG XL images could be made, but it seems that it is more of a theoretical question since real use is always relatively high BPP: humans are 1000x more expensive than computers, so human experience can be prioritized over computer working harder for us.
> only complicate coding at bitrates above 1.5 or so
these features could be disabled at higher bitrates as they're not helpful there?
But consider also e.g. https://afontenot.github.io/image-formats-comparison/#reykja... - even at "large" quality settings avif pretty visibly distorts the sky and there's some blur it the nearby trees too.
In any case, comparing to even the fairly good (for jpg) mozjpeg encoder it's clear both of these codes are much better than the status quo, and not that different from each other - neither wins universally vs each other, but both pretty clearly do vs. jpeg.
A fairly simple heuristic seems be that if want images at the tiny size - pick avif. At small, pick avif unless you really, really want to preserve texture over detail. At medium, pick jxl for texture, and avif for detail; and at large, pick jxl.
Browsing through these images in general, I think I'd usually pick jxl at medium or even sometimes large settings; small simply has too many artifacts in general (but if I had to use that - avif), and at better quality I (personally) find the distortion to texture more noticeable than loss of detail. I guess it depends on how important compression ratio is to you?
People in the internet don't like to store photos with 0.5 BPP even with the latest and greatest codecs, it gets too blurry and artefacty.
This is not a statement of my personal aesthetic opinion but observing what gets done out in the wild.
Usually we store images at 3.5 - 5 BPP (Cameras), 10+ BPP (Raw or similar for editing) and 1-2 BPP for internet user.
The actual bitrates depend a lot on the image -- graphics with simple backgrounds need less, photos with a lot of sky need less, busy detailed images particularly nature needs more.
While there is one 1.0 in the test, it is for an extremely busy image which would be better stored for internet use at 3+ BPP.
0.22 BPP is almost never used for photographs in the Internet.
See https://almanac.httparchive.org/en/2022/media#fig-15 for median bitrates in 2022 (1.0, 1.4 and 2.1 BPP for lossy formats)
> Usually we store images at 3.5 - 5 BPP (Cameras)
That's appropriate for JPEGs, but given that more recent compression algorithms do a better job, it's probably worth looking at lower bitrates. I pulled some JPEGs off my Canon DSLR for example, and they're around 2-4 bpp for landscape photos.
It's not surprising to see acceptable JPEG XL images with half that bitrate.
The final stage of their Image Quality QC was always "The Old Men With Good Eyes."
There were some employees that were professional "eyballers." They had full veto power.
I'm not sure if they still do that, but it would not surprise me, if they did.
[1]: http://dkriesel.com/en/blog/2013/0802_xerox-workcentres_are_...
Thanks for your good work!
What I like about JPEG XL, is the alpha-channel support.
When zooming in 3x I can say that AVIF looks significantly better than JPEGXL in the lower quality settings while being consistently a little bit smaller in filesize. There is better color retention, less noise and looks sharper. At the "tiny" setting the difference is night and day. At the highest setting they are so close enough that I wouldn't say there's a clear winner for me after going through all the images.
Disclaimer: JPEG XL contributor.
I agree that the example you presented indeed has the mentioned shadow issue in AVIF.
In reality cameras produce about 3.5-4.5 bpp, and the internet uses an average of 2 bpp for photographs.
Yep. It is ultimately a subjective experience. I found this out when implementing a color management system back in the early to mid 90s.
There is a lot of math, and a lot of physics, of the light, of boundary conditions of ink, etc. But there is also a lot of perception and psychology that is very difficult to capture.
As an example, my first attempts were quite accurate as measured, but looked horribly yellow-ish. The reason is that my software was successfully compensating for the blue-ish optical whiteners (UV+blue) in the paper. Which you can't do, because the eye will judge the brightest "nearly white" area in the field of view as white, and judge all the other colors relative to that.
But you also can't not do it, because then you ignore what the colors are actually supposed to be. And then it gets tricky...
Sometimes it's really clear. E.g. for the "US Open" Image, AVIF "tiny" (23.6 KB) looks as good as JPEG XL "medium" (46.3 KB), despite the substantial difference in bit rate:
https://afontenot.github.io/image-formats-comparison/#us-ope...
In some other cases it's less clear, and sometimes AVIF denoises too aggressively, e.g. in animal fur. JPEG XL has more problems with color bleed. Overall it seems to me AVIF is significantly better, especially on lower bit rates.
When doing a visual comparison, imo the best way is to start with the original versus a codec, to find out what bitrate you consider "good enough" for that image — this can vary wildly between images (on some images "small" will be fine while on others even "large" is not quite good enough). Then compare the various codecs at that bitrate.
There's a temptation to compare things at low qualities (e.g. "tiny") because there it's of course easier to see the artifacts. You cannot extrapolate codec performance at low quality settings to high quality though, e.g. it's not because AVIF looks better than JXL at "tiny" that it also looks better at "large". So if you want to do a meaningful comparison, it's best to compare at the quality you actually want to use.
At Cloudinary we did a large subjective study on 250 different images, at the quality range we consider relevant for web delivery (medium quality to near-visually lossless). We collected 1.4 million opinions via crowdsourcing in order to get accurate mean opinion scores. The results are available at https://cloudinary.com/labs/cid22.
One important thing to notice is that codec performance depends not only on the codec itself but also on the encoder and the encoder settings that are used. If you spend more time on encoding, you can get better results. A fair comparison is one that uses the best available encoders, at similar (and relevant) bitrates, and at similar (and relevant) encode speeds.
It's almost impossible to do subjective evaluation for all possible encoder settings though, or to redo the evaluations each time a new encoder version is released. This is why objective metrics are useful. There are many metrics, and some are better than others. You can measure how good a metric is by measuring how well it correlates with subjective results. According to our experiments, currently the best metrics are SSIMULACRA 2, Butteraugli 3-norm, and DSSIM. Older metrics like PSNR, SSIM, or even VMAF do not perform that well — probably indeed partially because some encoders are optimizing for them.
Here are some aggregated interactive plots that show both compression gains (percentage saved over unoptimized JPEG, at a given metric score) and encode speed (megapixels per second):
SSIMULACRA 2: https://sneyers.info/benchmarks/tradeoff-relative-SSIMULACRA...
Butteraugli: https://sneyers.info/benchmarks/tradeoff-relative-Butteraugl...
DSSIM: https://sneyers.info/benchmarks/tradeoff-relative-DSSIM.html
(Note that Butteraugli and DSSIM are distortion metrics, not quality metrics, so "less is better")