The case for JPEG XL
cloudinary.com
cloudinary.com
1) e-commerce textures such as cloths are more faitfully represented -- more trust in e-commerce, more revenue
2) more equity for selfies, people-of-color skin is better represented in JPEG XL, selfies in general look more people like with less smoothing of skin performed at compression stage -- I'm a great friend of the idea that filtering should not be built into the compression so that all kinds of images can be transmitted
3) more compression, faster and cheaper to load
4) better progression -- unlike I believe AVIF does, JPEG XL progression does not impose more decoding cost for sequential, progression does not make compression density worse
5) thermal camera support -- I believe thermal cameras have huge ecological and cost optimization opportunities and more phones should have thermal cameras in them to facilitate the shift away from fossil fuels
6) small images -- thumbnails and sprites etc. do not necessarily need spriting with JPEG XL. That is a nice increase in system level simplicity. JPEG XL overhead is about 10-20 bytes whereas AVIF overhead seems to be 600 bytes or so.
7) 'it just works, no surprises' -- the worst case image quality degradation in any quality setting are much less severe and much less common than in other formats
wait what? What is this woodoo and why it itsn't the case for other image formats in which encodings color in RGB?
> why it itsn't the case for other image formats in which encodings color in RGB?
The term "RGB" doesn't mean anything w.r.t. gamut, gamma, dynamic-range, nor how necessarily non-linear transformations between color-spaces are performed. But all that's irrelevant: JPEG'S colour support is actually quite fine (...if you don't mind chroma subsampling), whereas the problem here relates to how JPEG decides what part of the image are important for high-quality preservation (more bits, more quality) than other areas (less bits, less quality).
Obligatory ADHD relevant-yet-irrelevant diversion: this excellent (fun and interactive!) article about how JPEG works: https://parametric.press/issue-01/unraveling-the-jpeg/
(Caution: incoming egregious oversimplification of JPEG):
Now, just imagine a RAW/DNG photo of the sky with some well-defined cumulus clouds in, and another RAW/DNG photo of equal pixel dimensions and bit-depth (so if they were *.bmp files they'd have identical size), except the second photo shows only a grey overcast sky (i.e. the sky is just one giant miserable old grey duvet that needs washing). Now, the _information-theoretic_ content of the cumulus clouds photo is higher than the overcast clouds photo; the clearly-defined edges of the cumulus clouds are considered "high-frequency" data while the low-constrast bulbous gradients that make-up the overcast photo are considered "low-frequency" data, and less overall information is required to reconstruct or represent that overcast sky compared to the cumulus sky.
This relates to JPEG because JPEG isn't simply a single "JPEG algorithm", it's actually a sequence (or pipeline?) of completely different algorithms that each operate on the previous algorithm's output. The specific step that matters here is the part of JPEG that was designed with the knowledge that human vision perception doesn't need any high-frequency data in a scene where the subject is largely low-frequency data (i.e. we don't need a JPEG to preserve fine details in overcast clouds to correctly interpret a photo of overcast clouds, whereas contrarywise we do need JPEG to preserve fine-details that make-up the high-constrast edges between cumulus clouds and the blue sky behind them, otherwise we might think it's some other cloud formation with less well-defined edges (Altrostatus? Cirrostatus? Cirrus? I'm not a meteorologist - just someone on the spectrum using abusing analogies).
So my point so-far is that JPEG looks for and preserves the details in areas of images (where high-constrast data is) while disregarding high-frequency data in an overall low-frequency scene where it thinks those details don't matter to us humans with our weird-to-a-computer visual perception system. I imagine by now I've made JPEG sound like a very effective image compression algorithm (well, it is...) - so where's the problem?
...well, consider that when a person has a low-level of melanin in their skin they will feature areas of high-contrast on their faces: their skin's canvas is light - while facial-features like folds, creases, crows' legs, varicose veins, spots, etc, are like the hard-edges of those cumulous clouds from earlier: the original (pre-compression) image data will feature a relatively high image constrast where those features lie, and JPEG will try to preserve the details of that constrast by using more bits. So that's neat: JPEG tries to ensure that the parts of us that of us that are detailed remain detailed after compression.
...but what if a person's face is naturally less...contrast-y? That's the problem: when a person simply has darker skin their faces will feature less contrast between the skin-parts of their face compared to the features of their face - as caught by a camera. So if a simplistic JPEG compressor is being run against a United Colors of Benetton coffee-table book or stills taken from the end of a certain Michael Jackson music video that I'm particularly fond of[3] then we will, as you can imagine, unfortunately, see the JPEG image files of people with darker skin will feature less definition and detail of their faces compared to the photos of people with lighter-skin, simply due to how areas of higher, and lower, contrast are _mechanically_ processed by JPEG.
...which means that sloppily applying JPEG[1] to photos of everyone might work great for people who have faces-with-naturally-high-contrast/high-frequency-information on them - but not-so-great for everyone else - and I think we can agree hat we can improve on that.
-----------
Confounding things, the distribution of melanin in the JPEG development working group was not representative of the general population in the US at the time[5], let alone the rest of the world, which in-practice meant that less thought (if any) was paid to JPEG's stock handling of people with lower overall skin feature contrast (i.e. black people) - though no malice or ill-will or anything of the sort is required: like all of us (I assume...) they probably thought "these photographs of me (or people who generally look like me) processed by JPEG are good enough, it's been a long week, let's ship it and go for a pint" - without realising or appreciating that it wasn't "good enough" for large chunks of the population. At least, I hope that's the explanation... (and given Hanlon's Razor too).
The situation can be compared to the... uh... colourful stories about the history of Kodak's colour film, its shortcomings, and the social-consequences thereof[2]. Or for something more recent: the misbehaving face-detecting webcams of the last decade that lead HP to make a certain public statement[3].
-----------
[1] I say carelesslessly because lots of tooling for JPEG, especially at the pro-level (think: Photoshop) gives you a lot of control over how much information-loss is allowed for each macroblock (e.g. Photohop's Save-for-Web dialog (the old, good one, not the new dumbed-down one) lets you use its own paintbrush tool to mask areas that should lose fewer bits (i.e. higher-quality) even if the JPEG compressor doesn't think there's much value in that part of the image. MSPaint, on the other hand, does not.
[2] https://www.nytimes.com/2019/04/25/lens/sarah-lewis-racial-b...
[3] https://www.reuters.com/article/urnidgns852573c4006938800025...
[4] https://www.youtube.com/watch?v=F2AitTPI5U0
[5] I would like to mention that there was at least something for gender representation: the co-inventor of JPEG was a woman (no, I'm not referring to Lena.jpg) - Joan L. Mitchell, who sadly passed-away relatively young a few years ago: https://en.wikipedia.org/wiki/Joan_L._Mitchell
https://mobile.twitter.com/jonsneyers/status/155021585930558...
But I’m not clear if this was a design goal of JPEG XL, or a theoretical side-benefit, or something you would need an HDR screen to appreciate (I couldn’t tell the difference in any of the example images, but those may have been processed by Twitter).
> Concordia University professor Lorna Roth’s research has shown that it took complaints from corporate furniture and chocolate manufacturers in the 1960s and 1970s for Kodak to start to fix color photography’s bias.
.. but without excusing any of that, black people were still better off with the newer films.
Similarly, regardless of the intent of the JXL design, if it is the more "inclusive" tech then that speaks in favor of it.
Also, the difference in the example images requires you to look actively look for JPG artifacts. That's something that usually only codec developers and people who work with images for their day job do. Almost everyone else subconsciously does the opposite: they learn how to "look past" the compression artifacts. But speaking as someone who did spend a lot of time trying to optimize JPG sizes for a short while in his life there is definitely a difference.
The difference is hard to see when compressing an image once, but keep in mind that we live in a world where images go through lossy recompression dozens of times, causing the "needs more JPEG" meme[0][1] to live on even when it should be a thing of the past. If your image is, say, 5% more visually lossy for black people that adds up.
I can clearly see the differences between the encodings (I shuffled the images and sorted them by perceived quality), and I disagree that JXL looks better / is less racist. For both sides, the AVIF version looks significantly better, but especially for the black woman. The white woman already has fewer visible details that can be lost, the black woman has more, and JXL just smoothes those over, making it look like someone used the smudge tool (in general, but it’s especially bad here [0]).
So going from this example, I’d say AVIF wins by quite some margin.
I don't have the original decoded images nearby but it is best to do these kinds of comparisons without any recompression.
I have been able to (by both reducing the quality range and turning smoothing off) to get much better results on a high contrast image with most of its details in darker, having far-more details in those dark area with AVIF than JPEG-XL (not small differences either as you could see brick work that was destroyed by compression artifacts in JPEG-XL).
libavif default settings suck.
I admit the wording is kind of ambiguous though
- where is the best documentation for the Butteraugli metric? I’ve found it quite difficult to find a high-level rationale/design notes for the metric.
- I also understand that JXL doesn’t try to explicitly optimize against Butteraugli distance unless you invoke the higher effort levels, and then, it uses a slightly different function than the original Butteraugli distance - is that correct?
- One q about the distance parameter -d; if I pass -d 1 -E x for some x, I’m asking for “an image with at most Butteraugli distance 1”, and if I pick another effort level higher than x, am I asking it to introduce a little bit more loss, so that the Butteraugli distance is likely to increase? Ie is the distance parameter more of an “upper bound for acceptable loss”
- I’ve seen Butteraugli values quoted for different “norms/p-norms” - which one do I use?
- How do I learn more about these colorspace transfer functions?
Thanks again for your awesome work!
Butteraugli is largerly a 'learned' metric, just very small with its ~200 parameters. I implemented possible components and tuned them and their interactions using methods similar to machine learning.
Describing it to humans should be possible, but going all the way is going to be the same as describing DNNs.
SSIMULACRA 2 is taking the same approach and getting far better results than a previous effort SSIMULACRA.
2) butteraugli distance invoked in the higher effort levels
JXL heuristics are optimized to give good butteraugli scores. These heuristics are mostly used at speeds 6 and 7.
Speed 8 and 9 no longer trust solely on the heuristics, but they try to optimize a constant butteraugli score. It is important to note that they don't do RD-optimization, so they are quite bad at lower quality -- they will drive the quality bad no matter what. There is a hard limit for not reducing the quality too much in relation to the heuristics at speeds 6 and 7, but it is not RD optimization. Small savings can lead to large degradations. This is ok with near visually lossless (quality 90+), but not good for low quality (< 50).
Butteraugli has slightly evolved through the times, the latest version is available in libjxl and used by cjxl at effort 8 and 9.
3) -d 1 “an image with at most Butteraugli distance 1” ?
It used to work that way, but there was some inflation of this concept, and often the maximum score is twice the asked value. Ideally the maximum should be the given value, but it is very expensive to guarantee. In guetzli we have such guarantees, but it runs 1000x slower.
4) p-norm
usual metrics (SSIM, PSNR, VMAF) aggregate error over the image to a single number using the 2nd norm
we use a higher norm, and actually a combination of norms. If you ask for a 3rd norm, it is an equal combination of 3rd, 6th and 12th norm. Using higher norms puts more emphasis on large errors somewhere than 2nd norm.
The better the quality in image compression, best human rater correlations can be found in higher norms. 3rd norm is a good allrounder, I'd usually use 6th norm myself. A lower norm can be better for automated optimization since the cost function will have less wrinkles.
5) colorspace
we don't have a good write up other than possibly the JPEG XL spec. there is a javascript implementation of xyb here: https://gist.github.com/mattdesl/c0ad8d54b3491c7fd39e222e698...
Please consider joining the JPEG XL discord. I'm available for guiding you to learn more if you like.
Though I don't think thermal cameras will be user affordable for a long while.
There are a bunch of Android phones which already have this tech built in. Some are high end, but a surprising number of mid-range devices as well.
The ones that let you just walk into a house and immediately see cold areas cost more like $800.
For example, I remember reading somewhere that there are people looking at whether or not it makes sense to convert all of the existing JPEGs to JPEG XL on their servers for storage, then reconstruct the JPEG when needed. If so, then there is a niche where it can find adoption even if browsers don't support it yet. It also looks like it might find a use for medical and scientific imaging.
What I personally really would like to know is if it would be possible to losslessly recompress the RAW files from my camera. It looks like it should support everything required, and the claim that it "is the first serious candidate to become a universal image format that “works across the workflow”, in the sense that it is suitable for the lifecycle of a digital image, from capture and authoring to interchange, archival, and delivery" kind of implies that it is, but I haven't seen anyone say it out loud yet.
Because I must say: DNG is really slow and clunky to work with, and I bet has terrible compression on top of that.
This isn't a fight about whether something is technically superior, it is about marketing, cult and ideology. It wont be long before major force of AOM supporters arrives with new PR points to fight back.
Has there ever been a case when a company has demanded a device not support a codec that wasn't a patent dispute? And what does it mean for a device not to support an image format? Almost all image decoding is done in software.
This seems like an unfounded concern.
Ah, like how Chrome delayed APNG so much that there was a dedicated extension for it? (https://github.com/davidmz/apng-chrome) Sadly the Chrome extension page was removed around June (I assume it's more of cleaning up obsolete extensions rather than something nefarious) but I have recovered its description here:
Support for animated PNG images in Google Chrome browser.
Chrome 59 now have native support of APNG, so you do not need this extension anymore. Goodbye, it was an interesting 6 years!
About APNG format: http://en.wikipedia.org/wiki/APNG
This extension animates IMG elements. Also it animates background images/list style images from css styles (div style="…"), but CSS images support may be incomplete.
You can prevent this extension from work on specific domains (Black List mode, default) or allowed it to work ONLY on specific domains (White List mode).
Source code is available on GitHub: https://github.com/davidmz/apng-chrome
WARNING! CAPTCHA may not work on some sites because of this expansion (it is not always possible to correctly detect captcha images). If captcha does not work on some site, just (temporarily) disable the extension on the it.
---------------------------------- Latest notable changes:
3.1.0 (2017-06-07) Chrome 59 now have native support of APNG, so you do not need this extension anymore. Goodbye, it was an interesting 6 years!
3.0.0 (2016-09-12) IT WORKS AGAIN!!!
In Chrome 48 (Feb 2016) Google turns off getCSSCanvasContext API. It was the critical part of extension, and within six months the Chrome has no technologies to animate images. Fortunately, it has changed recently and now this extension works again.
In version 3 algorithm was changed from canvas animation to APNG → animated WebP conversion.
2.1.4 (2014-11-28) Fixed some incorrect captcha images.
2.1.2 (2014-01-15) Better Reddit support.
2.1.0 (2014-01-12) Partial Reddit emotes support.
2.0.7 (2014-01-12) Redirects support.
2.0.4 (2013-05-26) Ext. was broken because of bug 238071 in Chrome 27. Fixed it.
2.0.3 (2013-03-19) Fixed bug with images with 'auto' height/width in style.
2.0.0 (2012-10-01) Most of the code rewritten. CSS background/list animations is turned off (except inline styles). Fixed slow work in some sites (g+, gmail and so on).
0.7.1 (2012-04-07) animated-gif-captcha images support (strange captchas of wedge.org)
0.7.0 (2011-08-03) fix for facepunch.com (perhaps the only site where this extension is actually used:))
0.6.9 (2011-07-14) data: url support, fixes for google maps images
0.6.8 (2011-07-09) fixed bug with CAPTCHA images
0.6.7 (2011-07-03) support for list-style-images
0.6.6 (2011-06-30) tracking changes in image's 'src'
0.6.2 (2011-06-30) background-image animation for internal and external CSS
> Ah, like how Chrome delayed APNG so much that there was a dedicated extension for it?
No? APNG never caught on like described in the GP.
I wasn't aware of this and this is actually crazy, and freaking brilliant from a migration point of view. I wonder how that was achieved. I didn't care so much about JPEG XL, now thanks to Google derping I want it everywhere.
The final version still had brunsli-like features and we surfaced those as a JPEG recompression system.
Coincidentally today I was testing how much we can shrink our archived videos. We have lots of videos recorded with the Intel hardware encoder in H264, so not even the state of the art of H264. I figured: pop it into handbrake, change the codec to H265 (X.265), and boom 30%+ reduction. No, the H265 file got 3X larger. Well the default handbrake quality might be higher than the original video, and now it's going to some spend bits trying to encode the artifacts in the source video, while adding artifacts of its own. Sure you can find a happy medium where the quality won't be too much worse and you reduce the file size in most cases. But it sure would be nice to be able to losslessly recompress those videos!
And you highlight a crucial thing here: being able to do so without fiddling with parameters for hours to fine-tune it or even just finding something that works at all.
Motion JPEG XL would be like in the movies where they have Motion JPEG 2000, but ~35-40 % more dense. We could add usual delta frames without motion without compromising quality criteria. This would get us in the 0.3 BPP range.
4k at 30 Hz would be 3840 x 2160 x 24 x 0.3 should need about 60 mbps (~ 7.5 MB/s), still doable for home internet speeds and would be visually lossless, a better experience than home movie streaming is today. (free startup idea) :-)
So that's all impressive and useful and probably quite good, but my first reaction is that sinking feeling from being faced with unexpected complexity. Is the JXL format a nightmare internally? Is it a nightmare to use the library as an inexperienced app developer? Do you need to care about those unusual features if you just want to display an image? I expect that it's all fine, since what little I've heard has been good. If so, kudos to the designers because that's an ambitious feature list. It feels uncommon for a format of any kind to set out to be all things to all people and succeed.
The JPEG XL format is most certainly not a nightmare internally. Compare the spec lengths: 681 pages for https://aomediacodec.github.io/av1-spec/av1-spec.pdf, 101 for https://www.iso.org/standard/77977.html. Or code sizes: 6.6 MB uncompressed for libjxl (encoder+decoder), 9.5 MB for dav1d (decoder only) plus 34.2 MB for SVT-AV1 (encoder).
Features such as spot colors are mostly just a matter of understanding a wide range of requirements; this particular one just boils down to an extra image layer.
And here we immediately see one important difference between them: the first link is to a 681-page PDF, which opens immediately in my browser. The second link is to a place where one can buy access to what it says is a 101-page PDF, after paying more money than it would cost to buy a basic laptop.
No need to be imprecise, unless you are trying to push some agenda. The cost is CHF 198, which is around USD 195.
That said, I do hope that ISO will change its policy to put specs behind a paywall.
libjxl C++ implementation is modern high-performance multi-threaded SIMD code, with layering, animation, streaming and progressive features, it can look scary. With some less emphasis on performance it could be written rather elegantly.
Technically, JPEG XL may be a superior contestant. But from a 'social' point of view? Disastrous, especially considering it already has "JPEG" in it. (To be fair, its also quite young...)
I really hope JPEG XL gets more traction because it seems like a really good successor to the old namesake.
Maybe a nimbler PR approach would help, but then we know that standards like websql are shutdown over nothing, even though sqlite eating the world alive…
Anyway, I think "jxl" will in practice be the more common way to refer to JPEG XL images, and the full name will at some point become a matter of etymology and trivia questions.
And what could also be an advantage for JXL is animation. WEBP does have animation support implemented in browsers but it came after the freature was enabled and doesn't have a separate mime type so you can't do fallback to .gif without user agent checks. If JXL support either ships with animations enabled or animated JXL gets a separate mime type in browsers it could make it actually useful.
Particularly, WebP lossless does not use YUV compression. It does (optionally) reduce component correlations by the 'subtract green' transform.
It might be possible to construct an ICC profile in that formalism that would go close to YUV, but that is definitely not supported currently in cwebp.
Also, the ui doesn’t feel naive at all
Right click on text, there's no dictionary lookup menu item
The border around input fields look weird
Firefox has no search engine monopoly to advertise and build its brand upon.
Reasons why I, a former Firefox promoter switched to Chrome:
1. Speed. Chrome was blazing fast compared to other browser
2. Stability: Chrome could have a tab crash without bringing down your entire browser.
3. Really good developer tools that were superior to Firebug. As a matter of fact, Chrome Dev tools were my gateway drug as I'd use Chrome for work and Firefox for personal browsing, and the gradually shifted entirely to Chrome.
It is possible that FF won't continue the work of implementing it if they see no future without the oligarch in the field, but we'll see
All browser vendors are pretty reluctant to add new codecs, because there's always going to be yet another promising codec to add, but they're left maintaining all of them forever, even after they're not cool and new any more.
8-bits per channel and YUV420 only were caused by hurrying it.
With JPEG XL we have the opposing mistake. We had a great format already in 2017, and spent 5 years more for improving it further.
Firefox never had a monopoly, unless you still count Netscape Navigator in that. IE was big until 2006-08, started dropping off quickly after that, Firefox won a big share in that time, but at some point Chrome just chipped away marketshare for reasons stated above.
I think bundling was a big deal by Google, something Microsoft got punished for, but Google never.
Not all of us have left Firefox behind of course, but we are certainly not what we used to be.
> shouldn't we just fight for more diversity in browserland?
We should; and especially in browser-engine-land. WebKit/Blink are becoming a monopoly..
I didn't switch to Chrome. I switched to orher browsers. For me the turning point was when Firefox started deprecating its old GUI and copying more and more the Chrome GUI: no menu, hidden preferencies, messing with the window's titlebar and implementing its own (very small) scrollbars, messing with DNS, etc.
I use Firefox because Chrome's updater kept trying to sneak past my firewall and one day it succeeded. As soon as it did, I switched. **No.**
I don't like Firefox because of all its issues (including the memory leak), but at least it supports trackpad gestures and overscroll and all that good stuff that Chrome doesn't. And it doesn't keep trying to backdoor my computer using a stupid hard-to-disable updater. One single JSON file and it obeys for good.
Until now I used firefox on my home ubuntu box (intel nuc). Yesterday it started crashing, and crashed about 15 times. I downloaded chromium and it didn't yet crash.
Now, I'll probably move to it for home computer browsing, too. Even when I don't agree with their decision on not enabling JPEG XL by default. If there is a browser that has it on by default, then I naturally switch to that instead :-) -- but I'm not aware of such. Perhaps Brave?
Gemini Lake chip?
I used Firefox for a long time, starting way back when it was brand-new and IE was the leading browser. I continued using it because they were innovating heavily in the browser space. After they stopped innovating, I kept using it because although Chrome was impressive and fast, they kept nagging you about logging into Google. Now, based on the fact that Mozilla basically seems to just do whatever Chrome does these days, combined with the unnecessary and infuriating UX redesigns, Firefox has no obvious advantages and quite a few drawbacks.
I don't love the fact that Vivaldi is closed-source and based on Chromium but at least I feel like I'm in the driver's seat while I'm using it.
I'll repeat a comment I wrote here earlier:
Our camera hardware supports HDR. RAW image formats support HDR. Our display devices support HDR. Videos support HDR. Smartphones can now record in HDR.
All the hardware is there ready for the future. Yet due to (30 year old) software limitations we're still cutting off HDR data when saving our precious memories as JPEG. All whites become #ffffff. The white of paper. The color of Word's document background color.
The gamut of color has expanded and we need image standards to expand with it.
https://helpx.adobe.com/si/camera-raw/using/hdr-output.html
JPEG XL has HDR always built-in into its modeling, both SDR and HDR images are modeled exactly the same way and give the exactly same visual quality with the same encoding parameters. Compositing SDR and HDR materials does not need conversions between colorspaces -- everything is in absolute color XYB colorspace.
Other codecs (AVIF, HEVC) need a different set of encoding parameters for obtaining the same visual quality when used for SDR or HDR. Compositing SDR and HDR materials necessarily needs color conversions, since they will have different normalizations. Knowing which HDR quality setting corresponds which SDR setting will become additional cognitive load for the users.
> JPEG XL can do lossless image compression in a way that beats existing formats (in particular PNG) in all ways: it can be faster to encode, produces smaller files, and more...
Does "in all ways" include any decode (not encode) speed numbers? Such numbers are dependent on the test suite, and I might be holding it wrong, but on my measurements, JPEG XL lossless encodes 10x slower and decodes 12x slower than PNG (the libpng implementation), and decodes 30x slower than PNG (the wuffs implementation). JPEG XL does admittedly win on compression ratio (smaller files), though.
https://github.com/nigeltao/qoir/blob/main/doc/full_benchmar...
Knobs are also less relevant for lossless decoding, where there's only one correct output for any given input.
In any case, my code (https://github.com/nigeltao/qoir/blob/main/adapter/jxl_adapt...) copy/pastes examples/encode_oneshot.cc from the libjxl repository, plus an additional JxlEncoderSetFrameLossless call.
So I'm using whatever knob values that the libjxl example uses, i.e. the default knob values. If there are better knob values for lossless, let me know.
At one extreme we have fjxl, see e.g. https://twitter.com/LucaVersari3/status/1485971553892323333?.... This is an extremely fast lossless jxl encoder, which is about 100 times faster than PNG encoding (libpng) and compresses about 10% better.
At the other extreme, there are the slowest settings of the reference encoder libjxl, which are 1000 times slower than libpng but compresses a lot better.
As for decode speed: this depends on whether you consider single-threaded or multi-threaded performance. In single-threaded decode performance, it is hard to beat PNG since it is basically just gunzip. PNG is inherently sequential though (unless you apply trickery at encode time to make parallel decode possible), while JXL is by design decodable in parallel. We are anticipating that in the future, the number of cores available on typical hardware will only grow, so this difference will naturally lead to JXL becoming faster to decode in practice.
The current reference software libjxl is quite optimized for the lossy case, but still has some room for improvement for the lossless case — in particular it does not have fast paths yet for the common case where the full precision of 32-bit per sample is not required.
So in conclusion, I think it is safe to say that JPEG XL beats PNG in all ways.
If you look solely at encode speed, fjxl loses to fpnge (also listed in full_benchmarks.txt) so I'm not convinced yet that JPEG XL beats PNG in all ways.
I believe that Apple have an unofficial PNG extension that facilitates multi-threaded decoding: https://github.com/w3c/PNG-spec/issues/45
And yes, fpnge is slightly faster than fjxl but also compresses significantly worse. It is likely that you could make an even faster but slightly worse version of fjxl that would beat fpnge. But I think the speed of fjxl is already good enough in practice.
The problem with PNG extensions for multi-threading is that it only works if you control the encode side too — most existing encoders will not use such an extension. If existing deployments need to be replaced anyway, you can just as well use a new format altogether.
I had not. Thanks for the tip. Fjxl encoding speed is now about 20x faster than libpng, not 10x, but still mid-table. Decoding speed is unaffected.
https://github.com/nigeltao/qoir/commit/c2fec840
> If existing deployments need to be replaced anyway, you can just as well use a new format altogether.
There's still an operational difference between rolling out multi-threaded PNG versus multi-threaded $OTHER_FORMAT.
At the file format level, MT PNG can be designed to be backwards compatible. Older PNG decoders simply ignore the new non-critical chunk.
At the software level, you can roll out encoders that emit MT-capable PNGs without having to also roll out new decoders everywhere (or 99% everywhere, or 'on all major browsers'). You can roll out new decoders only partially, and those upgraded places enjoy the benefits, but you don't break the decoders you don't control.
Also, in terms of additional code size, upgrading your PNG library from version N to version N+1 is probably smaller than adding a whole new $OTHER_FORMAT library.
If you care about speed or transmission cost, then you try to look elsewhere.
WebP lossless already is quite a big improvement, and JPEG XL gives some more (not yet in decoding speed, though).
Anyway, yes, we should improve decode speed — for example, currently every sample value is unnecessarily converted from int to float and then back to int, and obviously that causes some avoidable slowdowns.
I agree that MT PNG does have a gentler transition path than introducing a new format. The point remains though that you need to somehow get both the encoder side and the decoder side upgraded to benefit from the advantage, which can be hard given how many existing deployments of png there are. Any system that has to take arbitrary png as input cannot rely on MT decode being an option, etc.
Also one disadvantage of MT PNG compared to JXL is that it splits the image in stripes, not tiles. For full-image decode, that doesn't matter much, but for region-of-interest (cropped) decode, tiles are a bit more efficient (no need to unnecessarily decode the areas to the left and right of the region of interest).
In case it wasn't clear, the 0.630 fjxl encode speed number in the top level README.md in that commit is after normalization such that QOIR is 1.000. After all, absolute numbers are hardware dependent.
If you look at the doc/full_benchmarks.txt change in the same commit I linked to, fjxl encode speed clocks at 106.48 megapixels per second. It's just that QOIR encode speed clocks at 168.90 megapixels per second (and fpnge at 312.59).
Well, for PNG, Apple have shipped exactly that. It was something they could unilaterally do without e.g. having to get all of the major browser makers on board.
Apple's PNG encoders (i.e iOS dev SDKs) produce multi-thread-friendly PNGs and their decoders (i.e. iOS devices) reap the benefits. And unlike trying to introduce a new but not-backwards-compatible-with-PNG format (JPEG-XL), other PNG decoders can still decode these images. They just don't get the speed benefit.
> Any system that has to take arbitrary png as input cannot rely on MT decode being an option, etc.
It's forwards compatible too. Apple's PNG decoder will use MT decode if the PNG image was encoded with MT metadata, but Apple's PNG decoder can still decode arbitrary PNG (using its pre-existing single-thread code path).
Also fjxl has an effort option, with 0 or 1 it will be even faster, but with less compression
Same for libjxl, there are faster efforts 1 and 2, and effort 3 as far as I know works well for photos but not so much for the rest
As for which is more important out of encode and decode speed, I would personally rank decode speed higher. For example, app icons on your home screen were encoded once (by the graphic artist) but are decoded zillions of times (by every user every day).
We have ideas how to make it faster than WebP lossless, but haven't engaged with that yet -- the lossless decoding is already fast enough not to be a blocker.
Also, given how efficient and quality-sure the lossy is, it will become tempting to migrate many of the PNG/GIF/WebP lossless cases to JPEG XL lossy further reducing the priority in JPEG XL lossless. This can be a 4x-5x savings whereas a better lossless is only 10-40 % savings in density.
I don't think having an existing library that all browsers could "just use" is a good argument. That just means that the implementation is the specification, no matter what the official spec says.
Note that I'm not saying that this is a problem or that JPEG XL shouldn't be supported. I just don't know what the real situation is, and this article smelled a little like it might be sweeping some things under the rug. (And for the other other side, the justification in the bug also seemed lacking.) Anyone know more?
You have a very good point. In fact I think this is the first time I ever seen such a question, which was also something I wanted to answer by making J40. Of course in the practical standpoint you should just use existing libraries, but I wanted to answer that from the library author's perspective.
I'm happy to report that JPEG XL's design is better than I hoped and it successfully tried to do many things out of a reasonable number of components. This means that library authors can mostly implement those components and they will combine reasonably well. There are still some imperfect edges though. For example JPEG XL splits a full image into multiple tiles to faciliate parallel decoding, and the Table of Contents is used to locate individual tile data. But TOC requires a full entropy coding to decode, so you can't easily determine how many bytes are needed for decoding, say, 1/2 of the full image. Those kind of imperfections are harmless for JPEG XL's main use cases but still something I would definitely fix if I'm even given a chance (not to say that I'm a good person to do so).
Google just doesn't like formats created outside their own company. They prefer WebP and AVIF (created by themselves), but dislike JPEG XL (not created by them).
This is a common misconception; JPEG XL is jointly created by Cloudinary (which employs the author of FLIP/FUIF) and Google Research Zurich (which previously created PIK, and employs authors of Brotli and WebP lossless format as well) with an official blessing by JPEG (which mostly set the goal for the format). The only difference here is that the authors do not belong to the Chrome team.
Apple is a member of AoM backing AV1/AVIF formats, together with other browser vendors.
No under Tim Cook though. They sided with Google already.
We’re still pretty committed to using jxl as an internal standard for many things.
I’m sure there are reasons making fresh 0.x releases even after creating a standard format, but it was a bit disappointing after years of waiting, even before this blow by certain browser maker.
That said, we are aiming to reach the libjxl 1.0 milestone within a reasonable timeframe, i.e. somewhere in 2023, preferably first half.
Perhaps unrelated but Google seems 'business tone deaf'. With every product release the first question on our minds is 'When will it get deprecated?' - and that is a legitimate question. Their support story is not that impressive. They do not seem to have the history/culture/DNA for support that MS or Amazon has. Their best customer is the browser - a customer that does not ask for support, only ads. I think they have prospered and will continue to for some time but they have to get a better non-browser customer support story.
Ad revenue paved a runway so wide that potholes do not bother them. If the revenues keep coming down they will need a shift in attitude (management and developers).
WebP's best quality is that in lossless mode, it usually results in smaller files than PNGs, but given that it was primarily focused at being a lossy format to replace JPEG, with a 16383×16383 size limit[1], color space fixed at YUV 4:2:0 only[2], marginal space savings, no progressive decoding, it comes up short in a number of scenarios. There's also the breadth of legacy JPEG-encoded files without a good transition story to WebP, especially due to most JPEGs in the world not being of YUV 4:2:0 format.
I sound like I'm hating WebP here and I'm not trying to, but my own personal uses for WebP have been squarely in the lossless mode side with a JPEG-encoded thumbnail to accompany it. WebP's primary selling point were file size benefits, but then mozjpeg really not just improved JPEG 1992, but made it a compelling choice to use instead of WebP.
JPEG XL really stands to be the best of everything. No fixed colorspace model, no image size limits, smaller lossy and lossless files than previous formats, progressive decoding, etc. It's a solid win for JPEG XL, especially if browsers start adopting it for production.
[1] Size limit applies to lossy and lossless mode both.
[2] Lossless mode's color space is fixed at 32-bit BGRA only. Same issue, different spec.
I remember all the same conversations happening now on HN, but for WebP back in the day, and Mozilla being really doubtful it was worth the maintenance effort. FF eventually added it... in 2019, and Safari in 2020. Only two years later we're discussing a format that will completely supplant it, but the browsers will have to maintain WebP forever.
Previously some engineers tried to bring bzip2 into content encoding, that failed by HTTP/tv-top-boxes interacting. HTTPS getting more common allowed a safer deployment route.
The W3C WG was trying to bring LZMA in, but that was just too slow for all fonts decoding -- it would have not made fonts load faster in all cases.
Not sure if it counts as momentum, but engineers representing Facebook, Intel/VESA, Shopify, Adobe, and more have been asking for JXL to be enabled in Chrome without a flag for quite a while now.
AFAICT, JPEG XL does not seem to have the licensing complications of JPEG2000.
Well, there is this thing https://jpegxl.io/articles/rans/
> Several variants of the coding procedure Asymmetric Numerical Systems (ANS) may be found in most modern codecs, such as AV1, Z-Standard compression, or even rANS in JPEG XL.
... this is hardly the fault of the JPEG XL developers though, that's the US patent system being ridiculous. Also, in the current discussion context that patent screws over any new compression format, not JPEG XL specifically. Note: general compression, not just image compression (ok, sure, QOI isn't affected. QOI also isn't relevant here).
It's obviously a defensive patent anyway, because between AV1, zstd and JPEG XL Microsoft would have to fight an alliance of literally everyone else out there if they'd plan on suing anyone.
simd based software coding if more than fast enough, at least for JPEG and JPEG XL
> At Cloudinary, we recently performed a large-scale subjective image quality assessment experiment, involving over 40,000 test subjects and 1.4 million scores.
This isn't about your, my or Jon Sneyer's individual opinions on these photos, it's the consensus of 40,000 test subjects. And yeah, I agree with you that for a significant number of photos I disagree with the consensus. But given that we're both posting on HN we're both also likely the kind of people who like to prove other people wrong by actively looking for and overvaluing counterexamples, so also keep that in mind. Because to be honest, for a large number of photos I also agree with the consensus.
As an example of the difficulties an academy leading corpora TID2013 has a reversal in quality between categories 4 and 5 -- highest quality images seem to get slightly worse quality ratings than the next highest quality images.
I observed a leading image quality laboratory to produce substantially different results for the same codec when the person conducting the experiments changed.
There are hard opinions if images should be reviewed one image pixel to one monitor pixel, or if images should be zoomed like they are in practical use, particularly on mobile.
Should zooming be allowed? Should people be allowed to move closer to the monitor if they want as part of the viewing? etc. etc.
I don't have time to dig through Twitter just now, but instead of just picking one image metric they've used all of them to also compare the testing methods themselves.
I guess scrollbar-gutter can fix the traditional scrollbar but then people have to give up using their fancy banners in headers. I'm (mostly) fine with Firefox's overlay scrollbar for now.
Mobile is a tricky tradeoff. There I am mostly used to Firefox, which has a tiny overlay scrollbar that is 50% grey making it almost invisibile in most cases.
It annoyed to me enough to make me activate overlay scrollbar on Firefox and inject custom CSS with scrollbar-gutter on Chromium. Otherwise, when moving from a repo home page (without a scroll bar) to the another page (let's say the contents a folder with a scrollbar), the box containing the contents of the repo experiences layout shift.
This is still an issue on GitHub and many other sites.
… and unfortunately it's not draggable, either.
Thank you.
[1] https://cloudinary.com/blog/time_for_next_gen_codecs_to_deth...
Instead visit https://storage.googleapis.com/demos.webmproject.org/webp/cm... and you will see that JPEG XL has color fringing and ringing artifacts.
> The source images were stripped of any metadata, EXIF, XMP, color profile etc. including the gAMA PNG chunk.
Just so happens that cjxl correctly supports PNG color profile/gamma data while cwebp doesn't.
Having said that, AVIF did get added and I don't think having some kind of process for adding new image formats beyond "it exists and is better" is a bad thing.
Ideally all the big players would come together and settle on one.
This kind of happened for AV1 video, so it's a shame that as a side effect of that AVIF got added before JPEG XL and split the potential momentum for a next generation image codec in two.
That feels like a mistake that it's not too late to revert though.
The Web platform is hardly cutting edge here. It took 10 years to add WebP across browsers, and it wouldn't have been added at all if it didn't cause web-compat issues for non-Chrome(ium) browsers.
From browsers' perspective the question isn't "is the new codec better", but "are existing codecs so terrible that they need urgent replacement". Most people still use JPEG, not even WebP or AVIF yet.
It sucks that the codec with the best quality/filesize ratio doesn't just win. There are many other factors beyond this, e.g.
• There's a lot of inertia keeping old codecs alive, which is from "just works" that has been from merely being old and popular. There's a whole graveyard of JPEG-killers that have beaten it on features and compression. GIF isn't even "good enough" technically, and it still refuses to die, because it's just so old to be universally supported. JPEG XL hasn't done the legwork yet to be everywhere, and isn't old enough yet to be everywhere including old software.
• Browser vendors' default resistance to ever-growing complexity of the platform, code size, and attack surface. I think JPEG XL's ability to be both non-Web editing format and Web format is actually working against it here, because the Web platform doesn't want non-Web features. Browsers haven't even implemented AVIF as specced, only a bare minimum subset. JPEG XL doesn't have a minimal subset.
So a variant of the "two for one" argument that justifies AVIF can also be used to justify JXL, since JPEG is needed anyway, even more so than AV1.
Also I am not so sure if the delta between AV1-only and AVIF is _that_ small, after all you do need to take care of alpha, color management, HDR, animation, and I suppose quite different code paths to get the decoded pixels in the right place than what is already there for the video case. Plus it looks like new things are still getting added to avif, like YCoCg-R and progressive previews, so it feels like it's a bit of a "moving target" spec like webp was in the early days. I see quite a few opportunities for bugs, different levels of support between different browser (versions), etc, that are not caused by AV1 itself but just by AVIF.
Level 5 (intended for browsers), allows no CMYK, limits the number of channels to 7, limits the maximum lossless bitdepth to 14-bit or so (the actual criterion is that the intermediate buffers can be implemented using int16_t, so the actual bitdepth limit depends on what bitdepth-widening or -narrowing transformations the encoder used), and has some more limits like that.
Level 10 allows nearly anything the bitstream syntax allows, with a few "sanity check" type limits.
All of the coding tools and format features in the jxl spec do have useful potential applications for the Web, except for CMYK and very high precision lossless (and this is why a decoder conforming only to level 5 does not need to handle those). E.g. layers can be useful just for compression too (e.g. a text overlay on a photo background will likely compress better / with less artifacts if it is kept in layers).
The nice thing about jxl is that everything needed to display an image is specced in the codestream and handled by libjxl itself, so no need for a browser (or any other application) to implement these things themselves — and possibly not implement some of them, or not quite correctly. In avif (like in heic), a lot of codestream gaps have to be filled at the file format level (e.g. alpha), and that inevitably leads to bugs and differences between the different browsers/viewers.
Given that, is there any reason JPEG XL might be better for privacy / worse for advertising industry scumbags? I have no expectation that Google / Chrome will ever make any decision primarily with their users' best interest in mind.
- JPEG XL: none
- HEIF: none
- AVIF: all but Edge
- WebP: all
https://www.cnet.com/tech/computing/chrome-banishes-photo-fo...
The JPEG-XL bug hasn't received any news or updates for at least a year now and the support is disabled on all stable builds (meaning: you can't even enable it with a flag since it's not compiled in). You can enable a flag on nightlies.
We can arrive to the best solution by listening, by creating and studying the evidence, and by discussion.
I'd like to highlight two things there:
My arguments for JPEG XL: https://github.com/mozilla/standards-positions/issues/522#is... -- it comes with links to 3rd party evidence of the somewhat outrageous but correct claims.
Jon's arguments for JPEG XL: https://github.com/mozilla/standards-positions/issues/522#is...
So I was wondering if/why it's not also part of Safari.
Google has a gigantic pile of money. They can afford to keep JPEG XL support.
Why is this manner of speaking so common?
This "we think supporting option X is too much work" is just corporate cost-cutting translated into developer-speak that is persuasive to those with no concept of history or user-centricity
People want options and they want diversity. How much maintenance is a JPEG XL decoder in the grand scheme of Chrome's codebase? Especially when it seems to be largely finished?
Although maybe I should have put "lazy-ass product managers" instead.
It's easy to feel powerless and frustrated in a world where corporations (in this case a blatant monopoly) make decisions that affect people's lives.
Without a second thought, because they arbitrarily deem it harmful to their profits or wasting their human resources. Good will... serving the public interest? Only if it's profitable or works towards the corporation's collective interests.
And it's not democratic: I can't propose or participate in a vote to overturn something that a corporation does or decides that affects my life. A Google employee couldn't even do that: Google is under no obligation to follow what their employees decide.
I can adapt, change my habits, even boycott Google; but ultimately my power as a citizen is limited because I lack the most important elements to affect change in this world: very big piles of money and influence.
It seems that Google's developers are merely cogs in the machine (see: the Stadia fiasco). So I'm sure the commenter was mainly directing their frustration to Google.
If Google's employees had a voice, surely Google would have a support department that handles Google's various services? How would you feel if countless people are despairing because the service you developed/worked on isn't serving them properly?
Locked out of your Gmail and recovery methods don't work? Google flagged your domain for so-called abuse? Take to Twitter if you're a celebrity or well-connected and beg for support, or if not: get fucked.
Google is not a good company, and desperately needs regulated.
This manner of speaking is common because not everybody is born with a silver spoon, or is able or willing to take on an enormous amount of debt to obtain a degree that may or may not pan out into a career that uplifts them out of poverty.
One interesting project would be a combined JPEG and JPEG XL decoder, written in rust which might even lead to a reduction in code size and better security since they have some overlap.
WebP appears to be written in C++, perhaps because underlying video codec is also written in C++…
[1] https://github.com/image-rs/image/tree/master/src/codecs/web...