JPEG XL and the Pareto Front
cloudinary.com
cloudinary.com
I've always thought that one as flying under the radar. Most get stuck on WebP not offering tangible enough benefits (or even worse) over MozJPEG encoding, but WebP _lossless_ is absolutely fantastic for performance/speed! PNG or even OptiPNG is far worse. And very well supported online now, and leaving the horrible lossless AVIF in the dust too of course.
Maybe you are thinking of high bit depth for archival use? I can see some use cases there where 8-bit is not sufficient, though personally I store high bit depth images in whatever raw format was produced by my camera (which is usually some variant of TIFF).
Also it's too power expensive to have your display in HDR if you're only showing SDR UIs, which not only is the common reality but will continue to be so for the foreseeable future.
Requiring every game, photo viewing app, drawing program, ... every application to decide how to do it's HDR->SDR seems unnecessary complication due to poor abstractions.
The locality of local tone mapping (the ideal approach to HDR->SDR mapping) would expose the window boundaries. Two photos or two halves of the same photo in different windows (as opposed to being in the same window) would create an artificial discontinuity for the correction fields being artificially contained within each window instead of spanning the users visual field as best as possible.
Every local tone mapping needs to make an assumption of the surrounding colors: is the window surrounded by black, gray, colored or bright light should influence how the tone mapping is done at borders. This information is not available for an app: it can only be done at the windowing system level or in the monitor.
The higher the quality of the HDR->SDR mapping in a system, the more opportunity there is to limit the maximum brightness, and thus also the opportunity for energy savings.
Games already do this regardless and it's part of their post processing pipelines that also govern aspects of their environmental look.
What's missing is HDR->Display which is why HDR games have you go through a clunky calibration process to attempt to reverse out what the display is going to the PQ signal, when what the game actually wants is what MacOS/iOS and now Android just give them - the exact amount of HDR headroom which the display doesn't manipulate.
As for your other examples, being able to target the display native range doesn't mean you have to. Every decent os has a compositor API that lets you offload this to the system.
https://zamundaaa.github.io/wayland/2023/12/18/update-on-hdr...
It is a good format when you can use it, but JPEG XL almost always compresses better anyway, and lacks color space and dimension limits.
The constraint was in WebP lossy to facilitate exact compatibility with VP8 specification and hoping that it would allow hardware decoding and encoding of WebP images using VP8 hardware.
Hardware encoding and decoding were never used, but the limitation stuck.
There was no serious plan to do hardware lossless, but the constraint was copied for "reducing confusion".
I didn't and don't like it that much as more PNG images couldn't be represented as WebP lossless as a result of that.
My desktop had a mixture of PNG and WebP files solely because of this limitation. I use the past tense because they've now all been converted to JPEG XL lossless.
HDR is (or can be) good for video & photography, but it's absolutely ass for UI.
Besides, you can just throw a gainmap approach at it if you really care. Works great with jpeg, and gainmaps are being added to heif & avif as well, no reason jpegxl couldn't get the same treatment. The lack of "true 10-bit" is significantly less impactful at that point
Gainmaps "solve" the problem of computing a local tone mapping by declaring that it needs to be done at server side or at image creation time rather than at viewing time.
My prediction: Gainmaps are going to be too complex of a solution for us as a community and we are going to find something else that is easier. Perhaps we could end up standardizing a small set of local tone mapping algorithms applied at viewing time.
Which was already the case. A huge amount of phone camera quality is from advanced processing, not sensor improveds. Trying to get that same level of investment in all downstream clients is both unrealistic and significantly harder. A big aspect of why DolbyVision looks better is just Dolby forces everyone to use their tone mapping, and client consistency is critical.
Gainmaps also avoid the proprietary metadata disaster that plaugues HLG/PQ video content.
> If anything you get more banding: two bandings, one banding from each of the two 8-bit fields multiplied together.
The math works out such that you get the equivalent of something like 9 bits of depth but you're also not wasting bits on colors and luminance ranges you aren't using like you are with bt2020 hlg or PQ
Mixing two independent quantization sources will lead to more error.
Some decoding systems such as traditional jpeg does not specify results exactly so bit-perfect quantization-aware compensation is not going to be possible.
Go try the math out, the theoretical is higher than 9 bits but 8.5-9 bit equivalent is very achievable in practice. With two lossless 8 bit sources this is absolutely adequate for current displays, especially since again you're not wasting bits on things like PQ's absurd range.
Will this be added to webp? Probably not, it seems like a dead format regardless. But 8 bit isn't the death knell for HDR support as is otherwise believed.
You just cannot reach the best quality with 8 bits, not in SDR, not in HDR, not with gainmap HDR. Sometimes you don't care for a particular use case, and then 8 bits becomes acceptable. Many use cases remain where degradation by a compression system is unacceptable or creates too many complications.
uhh, no it isn't?
And gainmaps suck, take lots of space, don't reduce banding. Even SDR needs 10-bit in a lot of situations to not have banding.
> uhh, no it isn't?
Find me a single example of a UI in HDR for the UI components, not for photos/videos.
> Even SDR needs 10-bit in a lot of situations to not have banding.
You've probably been looking at 10-bit HDR content on an 8-bit display panel anyway (true 10-bit displays don't exist in mobile yet, for example). 8-bits works fine with a bit of dithering
Yes, but the place where you want dithering is in the display, not in the image. Dithering doesn't work with compression because it's exactly the type of high frequency detail that compression removes. It's much better to have a 10 bit image which makes the information low frequency (and lets the compression do a better job since a 10 bit image will naturally be more continuous than an 8 bit image since there is less rounding error in the pixels), and let the display do the dithering at the end.
WebP also has a near-lossless encoding mode based on lossless WebP specification that is mostly unadvertised, but should be preferred over real lossless in almost every use case. Often you can half the size without additional visible loss.
Unfortunately, that option doesn't seem to be available in gif2webp (I mostly use WebP for GIF images - as animated AVIF support is poor on browsers and that has an impact in interoperability)
If you compress a whole manga, PNG (via oxipng, optipng is basically deprecated) is still the way to go.
Another something not mentioned in here is that lossless JPEG2000 can be surprisingly good and fast on photographic content.
If I had just added 128 to the residuals, all remaining prediction arithmetic would have worked better and it would have given 1 % more density.
This is because most related arithmetic for predicting pixels is done in unsigned 8 bit arithmetic. Subtract green moves such predictions to often cross the 0 -> 255 boundary, and then averaging, deltas etc make little sense and add to the entropy.
- OxiPNG - 730k
- webp lossless max effort - 702k
- avif lossless max effort - 2.54MB (yay!)
- jpegxl lossless max effort - 506k (winner!)
Possibly mostly focused on medium and low quality.
The heuristics (here https://github.com/libjxl/libjxl/blob/main/lib/jxl/enc_ar_co...) for choosing that value are quite primitive and produces only two values: no smoothing and some smoothing (values 0 and 4).
If we replace those heuristics with a search that tries out which of the values is closest to the original, we should get better quality, especially at the lowest bitrates where smoothing is important.
Whereas jxl and avif just become blurry.
These images attempt to be at equal level of distortion, not at equal compression.
Bpps are reported beside the images.
In practice, use of quality 65 is rare in the internet and only used at the lowest quality tier sites. Quality 75 seems to be usual poor quality and quality 85 the average. I use quality 94 yuv444 or better when I need to compress.
Bitrates are in the left column, jpg low quality is the same size as jxl/avif med-low quality (0.4bpp), so you should compare the bottom left picture to the top mid and right pictures.
The author of the blog post did exactly that in a previous blog post:
https://cloudinary.com/labs/cid22/plots
Human ratings are expensive and clumsy so people often use computed aka objective metrics, too.
The best OSS metrics today are butteraugli, dssim and simulacra. The author is using one of them. None of the codecs was optimized for that metrics except jpegli partially.
Don't get swept away by false comparisons, JXL and AVIF look significantly better if you give them twice as much filesize to work with as well.
> Decode speed is not really a significant problem on modern computers, but it is interesting to take a quick look at the numbers.
There was a much more easy to notice impact from progressive images and even sequential images displayed in a streaming manner during the download. As a rule of thumb, sequential top-to-bottom streaming feels 2x faster as a waiting rendering, and progressive feels 2x faster than sequential streaming.
If a new format drains battery twice as fast, users don't want it.
But in any case, there are no _major_ differences in decoding speed between the various image formats. The difference caused by reducing the transfer size (network activity) and loading time (user looking at a blank screen while the image loads) is more important for battery life than the decoding speed itself. Also the difference between streaming/progressive decoding and non-streaming decoding probably has more impact than the decode speed itself, at least in the common scenario where the image is being loaded over a network.
For image gallery use of camera resolution photographs (12-50 Mpixels) it can be more fun to have 100+ Mpixels/s, even 300 Pixels/s.
OTOH video decoding is highly likely to be hardware accelerated on both laptops and smartphones.
> For still images, the main thing that drains your battery is the display, not the image decoding :)
I wonder if it becomes noticeable on image-heavy sites like tumblr, 500px, etc.
Mobile phone CPU can switch between different power state very quickly. If the image decoding is fast, it can sleep more
Very few applications are constantly decoding images. Today a single image is often decided in a few milliseconds, but watched 1000x longer. If you 10x or even 100x energy consumption of image decoding, it is still not going to compete with display, radio and video decoding as a battery drain.
But the person encoding is picking the format, not the decoder.
I don't believe QOI will ever have any sort of real-world practical use, but that's quite OK and I love it for it has made me and plenty of others look into binary file formats and compression and demystify it, and look further into it. I wrote a fully functional streaming codec for QOI, and it has taught me many things, and started me on other projects, either working with more complex file formats or thinking about how to improve upon QOI. I would probably never have gotten to this point if I tried the same thing starting with any other format, as they are at least an order of magnitude more complex, even for the simple ones.
Actually, there was a big push to add QOI to stuff a few years ago, specifically due to it being "fast". It was claimed that while it has worse compression, the speed can make it a worthy trade off.
Prusa (the 3d printer maker) seems to think otherwise! https://github.com/prusa3d/Prusa-Firmware-Buddy/releases/tag...
Also, curious that they only benchmarked QOI for "non-photographic images (manga)", where QOI fares quite badly because it doesn't have palleted mode. QOI does much better with photos.
> Not shown on the chart is QOI, which clocked in at 154 Mpx/s to achieve 17 bpp, which may be “quite OK” but is quite far from Pareto-optimal, considering the lowest effort setting of libjxl compresses down to 11.5 bpp at 427 Mpx/s (so it is 2.7 times as fast and the result is 32.5% smaller).
17 bpp is way outside the area shown in the graph. All the other results would've gotten squished and been harder to read, had QOI been shown.
I just ran qoibench on the photos they used[1] and QOI does indeed fair pretty badly with a compression ratio of 71.1% vs. 49.3% for PNG.
The photos in the QOI benchmark suite[2] somehow compress a lot better (e.g. photo_kodak/, photo_tecnick/ and photo_wikipedia/). I guess it's the film grain with the high resolution photos used in [1].
Not something a consumer knowingly uses, but also not quite irrelevant either.
bz2 is obsolete. It’s very slow, and not that good at compressing. zstd and lzma beat it on both compression and speed at the same time.
QOI’s only selling point is simplicity of implementation that doesn’t require a complex decompressor. Addition of bz2 completely defeats that. QOI’s poorly compressed data inside another compressor may even make overall compression worse. It could heve been a raw bitmap or a PNG with gzip replaced with zstd.
On the other hand, libvpx has always been a mediocre encoder which I think might be the reason for disappointing performance (I mean in general, not just speed) of vp8/vp9 formats, which inevitably also affected performance of lossy WebP. Dark Shikari even did a comparison of still image performances of x264 vs vp8 [0].
[0] https://web.archive.org/web/20150419071902/http://x264dev.mu...
scoop install main/libjxl
Note.. now that I tried it: that is really next level for an old format..!
That sounds like it might negatively affect compatibility with older decoders that do not handle color spaces correctly.
I believe Jon has compared jpegli without XYB. If you turn XYB on, you get about 10 % more compression.
Jpegli is great even without XYB. It has many other methods for success (largely copied over from JPEG XL adaptive quantization heuristics, more precise intermediate calculations, as well as the guetzli method for variable dead-zone quantization).
disclaimer: I created the XYB colorspace, most of the JPEG XL VarDCT quality-affecting heuristics, and scoped jpegli. Zoltan (from WOFF2/Brotli fame!) did the actual implementation and made it work so well.
We kept a lot of focus on visually lossless and I didn't want to add format features which would add complexity but not help at high quality settings.
In addition to modeling features, the context modeling and efficiency of entropy coding is critical at high quality. I consider AVIFs entropy coding ill-suited for high quality or lossless photography.
It's also mentioned in [1], which starts off
> Today we're sharing open source code that can sort arrays of numbers about ten times as fast as the C++ std::sort, and outperforms state of the art architecture-specific algorithms, while being portable across all modern CPU architectures. Below we discuss how we achieved this.
[0] https://github.com/google/highway
[1] https://opensource.googleblog.com/2022/06/Vectorized%20and%2..., which has an associated paper at https://arxiv.org/pdf/2205.05982.pdf.
PS: I used to work on JPEG XL. It is great to see these outstanding improvements, congrats to the team!
.. $ ls -l a.jpg && shasum a.jpg
... 615504 ... a.jpg
716744d950ecf9e5757c565041143775a810e10f a.jpg
.. $ cjxl a.jpg a.jxl
Read JPEG image with 615504 bytes.
Compressed to 537339 bytes including container
.. $ ls -l a.jxl
... 537339 ... a.jxl
But, wait for it: .. $ djxl a.jxl b.jpg
Read 537339 compressed bytes.
Reconstructed to JPEG.
.. $ ls -l b.jpg && shasum b.jpg
... 615504 ... b.jpg
716744d950ecf9e5757c565041143775a810e10f b.jpg
Do you realize how many billions of JPEG files there are out there which people want to keep? If you recompress your old JPEG files using a lossy format, you lower its quality.But with JPEG XL, you can save 15% to 30% and still, if you want, get your original JPG 100% identical, bit for bit.
That's wonderful.
P.S: I'm sadly on Debian stable (12 / Bookworm) which is on ImageMagick 6.9 and my Emacs uses (AFAIK) ImageMagick to display pictures. And JPEG XL support was only added in ImageMagick 7. I haven't looked more into that yet.
Very impressive! The article too is well written. Great work all around.
Does fast rav1e look better than jpegli at high encode speeds?
Edit: found https://github.com/xiph/rav1e/issues/2759
I’ve only compared rav1e to mozjpeg and libwebp, and at fastest speeds it’s only barely ahead.
https://www.youtube.com/watch?v=UphN1_7nP8U
https://www.youtube.com/watch?v=inQxEBn831w
There's more on the same channel, generation loss ones are really interesting.
Seriously, when is the last time mobile phones used hardware decoding for showing images? Flip phones in 2005?
I know camera apps use hardware encoding but I doubt gallery apps or browsers bother with going through the hardware decoding pipeline for hundreds of JPEG images you scroll through in seconds. And when it comes to showing a single image they'll still opt to software decoding because it's more flexible when it comes to integration, implementation, customization and format limits. So not surprisingly I'm not convinced when I repeatedly see this claim that mobile phones commonly use hardware decoding for image formats and software decoding speed doesn't matter.
If you don't really care about your users' battery life you can opt to disable hardware acceleration within your applications, but it's usually enabled by default, and for good reason.
I keep hearing and hearing this but nobody has ever yet provided a concrete real world example of smart phones using hw decoding for displaying images.
At best, the camera processors output encoded JPEG/HEIF for taken pictures, but that's about it.
The SoCs aren't investing more than a token amount of effort into those jpeg decoders, and from experience some of them claim to exist but produce the shittiest looking output imaginable and more slowly than jpeg-turbo at that.
Also you can trivially find out if your Android phone is doing this or not, just run some perf call sampling while decoding jpegs. If all you see is AOSP libraries & libjpeg-turbo, well then they aren't doing hardware decodes :)
And yes, regular JPEG is still a fine format. That's part of the point of the article. But for many use cases, better compression is always welcome. Also having features like alpha transparency, lossless, HDR etc can be quite desirable, and those things are not really possible in JPEG.
While in practice it won't change my life much, I like the elegance of using a modern standard with this level of performance an efficiency.
Image codecs are used in a wide range of attacker-controlled scenarios and need to be completely safe.
I know Rust advocates sound like a broken record, but this is the poster child for a library that should never have been even started in C++ in the first place.
It’s absolute insanity that we write codecs — pure functions — in an unsafe language that has a compiler that defaults to “anything goes” as an optimisation technique.
[1] https://github.com/tirr-c/jxl-oxide/pull/267
> It’s absolute insanity that we write codecs — pure functions — in an unsafe language that has a compiler that defaults to “anything goes” as an optimisation technique.
Rust and C++ are exactly the same in how they optimize, compilers for both assume that your code has zero UB. The difference is that Rust makes it much harder to accidentally have UB.
There are only a handful of image codecs that are widely accepted. Essentially just GIF, PNG, and JPG. There's a smattering of support for more modern formats, but those three dominate.
Adding a fourth image format is increasing this attack surface by a substantial margin across a huge range of software. Not just web browsers, but chat apps, server software (thumbnail generators), editors, etc...
This is the kind of thing that gets baked into standard libraries, operating systems, and frameworks. It's up there with JSON or XML.
You had better be damned sure what you're doing is not going to cause a long list of CVEs!
JPEG XL is a complex codec, with a lot of code. This increases the chance of bugs and increases the attack surface.
A (surprisingly!) good metric for complexity is the size of the zip file of the code. Libjpeg is something like 360 kB, libpng is 350 kB, and giflib is 90 kB.
The JXL source is 1.4 MB zipped, making more than twice the size of the above three combined!
The other libraries use C/C++ not because that's a better choice, but because it was the only choice back in the ... checks Wikipedia ... 1980s and 90s!
We live in the future. We have memory-safe languages now. We're allowed to use them. You won't get in trouble from anyone, I promise.
> We live in the future. We have memory-safe languages now. We're allowed to use them. You won't get in trouble from anyone, I promise.
That's why I specifically said that it's unfortunate that C++ is still wide spread, and pointed to a fully conformant JXL decoder written in Rust :p
> There are only a handful of image codecs that are widely accepted. Essentially just GIF, PNG, and JPG. There's a smattering of support for more modern formats, but those three dominate.
Every browser ships libwebp and an AVIF decoder. Every reasonably recent Android phone does as well. And every iPhone. Every (regular) install of Windows has libwebp. Every Mac has libwebp and dav1d. That's all C++. AVIF in particular is only a couple of years older than JXL, and yet I've never seen opposition to it on the grounds of memory safety. That is what I meant about JXL being singled out.
> JPEG XL is a complex codec, with a lot of code. This increases the chance of bugs and increases the attack surface.
> A (surprisingly!) good metric for complexity is the size of the zip file of the code. Libjpeg is something like 360 kB, libpng is 350 kB, and giflib is 90 kB.
> The JXL source is 1.4 MB zipped, making it nearly twice the size than all of the above combined.
Which code exactly are you including in that? The libjxl repo has a lot of stuff in it, including an entire brand new JPEG encoder! Though jxl certainly is more complex than those three combined, since JXL is essentially a superset of all their functionality, plus new stuff.
I also recompressed all of the libraries with identical settings to make the numbers more consistent.
No such thing.
> a compiler that defaults to “anything goes” as an optimisation technique
That's just FUD.
I have forever associated Webp with macroblocky, poor colors, and a general ungraceful degradation that doesn't really happen the same way even with old JPEG.
I am gonna go look at the complexity of the JXL decoder vs WebP. Curious if it's even practical to decode on embedded. JPEG is easily decodable, and you can do it in small pieces at a time to work within memory constraints.
That's improved somewhat, but the formats that will have an easy time winning are the ones that people can use, even if that means a browser should "save JPGXL as JPEG" for awhile or something.
There is pngquant:
> a command-line utility and a library for lossy compression of PNG images.
Mobile browsers seem to default to downloading in png as well.
No, JPEG XL files can't be viewed/decoded by software or devices that don't have a JPEG XL decoder.
Yeah, this not means what usually we call backwards compatibility, but allows usage like storing the images as JPEG XL and, on the fly, send a JPEG to clients that can't use it, without any loss of information. WebP can't do that.
And in general, Jon's posts provide a pretty good overview on the topic of codec comparison
Pity such a great format is being held back by the much less rigorous reviews
Also, the frontier isn't convex, so it's unlikely that if intermediate options could be added then they would all be at least as good as the lines shown; and the use of log(speed) for the y-axis affects what a straight line on the graph means. It's fine for giving a good view of the dataset, but if you're going to make a guess about intermediate possibilities, 'speed' or 'time' should also be considered.
Some of the intermediate options are available though, through various more fine-grained encoder settings than what is exposed via the overall effort setting. Of course they will not fall exactly on the line that was drawn, but as a first approximation, the line is probably closer to the truth than the staircase, which would be an underestimate of what can be done.
*
|
|
|
*-----+
or +-----*
|
|
|
*
… and why?Such a shame arithmetic coding (which is already in the standard) isn't widely supported in the real world. Because converting Huffman coded images losslessly to arithmetic coding provides an easy 5-10% size advantage in my tests.
Alien technology from the future indeed.
JPEG XL, Brunsli, Draco 3D and ZStd use table-based arithmetic coding (ANS).
JPEG, Deflate, Brotli, MP3, WebP lossless and the fastest mode of JPEG XL lossless use prefix coding.
One thing I think would help with its adoption, is if they would work with e.g. the libvips team to better implement it.
For example, streaming encoder and streaming decoder would be the preferred integration method in libvips.
But the King remains HALIC. In terms of MT encoder it still uses 3.5x more memory than HALIC, and 6x encoding time compared to HALIC. While offering the same or smaller files size in Lossless. Hopefully JPEG XL could narrow those gaps some days.
There are a lot of little considerations like this, and it would be well if the industry consolidated around an animated-image standard, one which was an image, and not a video embedded in a way which looks like an image.
You mean exponential golomb style, where numbers like 0b110100101 would be encoded as 00000000110100110 (essentially using 2 bits per bit)?
It's because WebP has a special encoding pipeline for lossless pictures (just like PNG) while AVIF is basically just asking a lossy encoder originally designed for video content to stop losing detail. Since it's not designed for that it's terrible for the job, taking lots of time and resources to produce a worse result.
I mean it's kinda hard to be worse than just shoving a bitmap through zlib...
https://issues.chromium.org/issues/40270698
https://bugs.chromium.org/p/chromium/issues/detail?id=145180...
Also, all those Chrome offshoots (Edge, Brave, Opera, etc) could easily add and enable it to distinguish themselves from Chrome ("faster page load", "less network use") and don't. Makes me wonder what's going on...
And I suppose the Chrome folks have the telemetry to know how many people set that damn flag.
> “On display? I eventually had to go down to the cellar to find them.”
> “That’s the display department.”
> “With a flashlight.”
> “Ah, well, the lights had probably gone.”
> “So had the stairs.”
> “But look, you found the notice, didn’t you?”
> “Yes,” said Arthur, “yes I did. It was on display in the bottom of a locked filing cabinet stuck in a disused lavatory with a sign on the door saying ‘Beware of the Leopard.’”
But "implement something new!" is a very different demand from "you took that away from us, undo that!"
Is asking for the old thing to be re-added, but without the flag that sabotaged it. It is the same as "you took that away from us, undo that!" Removing a flag does not turn it into a magical, mystical new thing that has to be built from scratch. This is silly. The entire point of having flags is to provide a testing platform for code that may one day have the flag removed.
That said: these actual users didn't demonstrate any hacker spirit or interest in using JXL in situations where they could. Where's the wide-spread use of jxl.js (https://github.com/niutech/jxl.js) to demonstrate that there are actual users desperate for native codec support? (aside: jxl.js is based on Squoosh, which is a product of GoogleChromeLabs) If JXL is sooo important, surely people would use whatever workaround they can employ, no matter if that convinces the Chrome team or not, simply because they benefit from using it, no?
Instead all I see is people _not_ exercising their freedom and initiative to support that best-thing-since-slices-bread-apparently format but whining that Chrome is oh-so-dominant and forces their choices of codecs upon everybody else.
Okay then...
In the current scenario jpeg xl users are most likely to emerge outside of the web, in professional and prosumer photography, and then we will have — unnecessarily — two different format worlds. Jpeg xl for photography processing and a variety of web formats, each with their problems.
How is that relevant? Flags are to allow testing, not to gauge interest from regular users.
Chrome has a neat feature where some flags can be enabled by websites, so that websites can choose to cooperate in testing. They never did this for JXL, but if they re-added JXL behind a flag, they could do so but with such testing enabled. Then they could get real data from websites actually using it, without committing to supporting it if it isn't useful.
> Also, all those Chrome offshoots (Edge, Brave, Opera, etc) could easily add and enable it to distinguish themselves from Chrome ("faster page load", "less network use") and don't. Makes me wonder what's going on...
Edge doesn't use Chrome's own codec support. It uses Windows's media framework. JXL is being added to it next year.
Interesting!
I recently come across another issue pertaining to the chromium team not budging on their decisions, despite pressure from the community and an RFC backing it up - in my case custom headers in WebSocket handshakes, that are supported by other Javascript runtimes like node and bun, but the chromium maintainer just disagrees with it - https://github.com/whatwg/websockets/issues/16#issuecomment-...
That hammer is very close to going away; if the EU does force Apple to really open the browsers on the iPhone, everything will be Chrome as far as the eye can see in short order. And then we fully enter the chromE6 phase.
Unless it is some kind of anti-competitive behavior like they intentionally stiffening adoption of standard competing with their proprietary patent-encumbered implementation that they expect to collect royalties for (doesn't seem to be the case), then I don't see the problem.
Firefox is "neutral", which I understand as meaning they'll do whatever Chrome does.
All the code has been written, patches to add JPEG XL support to Firefox and Chromium are available and some of the forks (Waterfox, Pale Moon, Thorium, Cromite) do have JPEG XL support.
https://github.com/niutech/jxl.js is based on Chromium tech (Squoosh from GoogleChromeLabs) and provides an opportunity to use JXL with no practical way for Chromium folks to intervene.
Even if that's a suboptimal solution, JXL's benefits supposedly should outweight the cost of integrating that, and yet I haven't seen actual JXL users running to that in droves.
So JXL might not be a good support for your theory: where people could do they still don't. Maybe the format isn't actually that important, it's just a popular meme to rehash.
Fwiw I agree that there's a weird narrative around jpegxl, at the end of the day it's just a format, and I think it's not very good for lower quality images as proven by the linked article in the OP. Avif looks better in that regard.
I think it would've made more sense than WebP though (which also doesn't look good at all when not lossless), but that was like a decade ago and that ship has sailed. So avif fills a niche that WebP sucks at, while jpegxl doesn't really do that. That alone is reason enough to not bother with including it.
Average/median quality of images is between 85 to 90 depending how you calculate it.
There, users' waiting time is worth during image formats life time for about 3 trillion USD. If we can reduce 20 % of it we create wealth of 600 billion USD distributed to the users. More savings come from data transfer costs.
I'm not assuming that there are those benefits, but that there are people to see them. Those who _very_ vocal about browsers (and Chrome in particular) not supporting it seem to think so or they wouldn't bother.
If I propose integrating good old Targa file support into Chrome, I'd also be asked about a cost/benefit analysis. And by building and using a polyfill to add that support, I show that I'm serious about Targa files, which gives credence to my cost/benefit analysis and also lets people play around with the Targa format, hopefully making it self-evident that the format is good, and from there that these benefits based on native support would be even better.
For JXL I see people talking the talk but, by and large, not walking the walk.
To fix this, you'd need to convince Google, and other large companies that would be exposed to law suits related to these patents (Apple, Adobe, etc.), that these patent holders are not going to insist on being compensated.
Other formats are less risky; especially the older ones. Jpeg is fine because it's been out there for so long that any patents applicable to it have long expired. Same with GIF, which once was held up by patents. Png is at this point also fine. If any patents applied at all they will soon have expired as the PNG standard dates back to 1997 and work on it depended on research from the seventies and eighties.
Adobe included JPEG XL support to their products and also the DNG specification. So that argument is pretty much dead, no?
Harder to do for users of Chrome.
https://helpx.adobe.com/content/dam/help/en/camera-raw/digit...
Do you have source for that claim?
I think it would be much better for everyone involved and humanity if Mr. Duda himself got the patent in the first place instead of praying no one else will.
And nothing advances your career quite like getting your employer into a multi-year legal battle and spending a few million on legal fees, to make some images 20% smaller and 100% less compatible.
A few years earlier, Google was granted a patent for ANS in general, which made people very angry. Fortunately they never did anything with it.
...
The fact that Google does have a patent which covers JXL is worrying though. So JXL is patent encumbered after all.
Apple has implemented JPEG XL support in macOS and iOS. Adobe has also implemented support for JPEG XL in their products.
Also, if patents were the reason Google removed JXL from Chrome, why would they make up technical reasons for doing so?
Please don't present unsourced conspiracy theories as if they were confirmed facts.
It is slower for decoding and Jpeg xl does not do that for decoding speed reasons.
The specification doesn't allow it. All coding tables need to be in final form.
When the big boys want to do something, they find a way to get it done, patents or no, especially if there's only "fear of patents" - see Apple and the whole watch fiasco.
https://bugzilla.mozilla.org/show_bug.cgi?id=1539075
It's a real shame, because this is one of those few areas where Firefox could have lead the charge instead of following in Chrome's footsteps. I remember when they first added APNG support and it took Chrome years to catch up, but I guess those days are gone.
Oddly enough, Safari is the only major browser that currently supports it despite regularly falling behind on tons of other cutting-edge web standards.
Ten years ago Mozilla used to have the most prominent image and video compression effort called Daala. They posted inspiring blog posts about their experiments. Some of their work was integrated with Cisco's Thor and On2's/Chrome's VP8/9/10, leading to AV1 and AVIF. Today, I believe, Mozilla has focused away from this research and the ex-Daala researchers have found new roles.
I am not sure I would say that is true.
The entire entropy coder, used by every tool, came from Daala (with changes in collaboration with others to reduce hardware complexity), as did some major tools like Chroma from Luma and the Constrained Directional Enhancement Filter (a merger of Daala's deringing and Thor's CLPF). There were also plenty of other improvements from the Daala team, such as structural things like pulling the entropy coder and other inter-frame state from reference frames instead of abstract "slots" like VP9 (important in real-time contexts where you can lose frames and not know what slots they would have updated) or better spatial prediction and coding for segment indices (important for block-level quantizer adjustments for better visual tuning). And that does not even touch on all of the contributions from other AOM members (scalable coding, the entire high-level syntax...).
Were there other things I wish we could have gotten in? Absolutely. But "done" is a feature.
The primary benefit of PVQ is the side-information-free activity masking. That is the sort of thing that cannot be judged via PSNR and requires careful subjective testing with human viewers. Not something you want to be rushing at the last minute. After gauging the rest of AOM's enthusiasm for the work, we decided instead to improve the existing segmentation coding to make it easier for encoders to do visual tuning after standardization. That was a much simpler task with much less risk, and it was adopted relatively easily. I still think it was the right call.
[1] https://datatracker.ietf.org/doc/html/draft-cho-netvc-applyp...
They probably gave up because they simply don’t have the money/resources to pursue this.