JPEG XL: How it started, how it’s going
cloudinary.com
cloudinary.com
Is this because almost all AV decoders use libffmpeg or a fork thereof; where libffmpeg is basically an "uber-library" that supports all interesting AV formats and codecs; and therefore you can expect ~everything to get support for a new codec whenever libffmpeg includes it (rather than some programs just never ending up supporting the codec)?
If so — is there a reason that there isn't a libffmpeg-like uber-library for image formats+codecs?
Sure, but I don't mean general-purpose mulimedia containers (that put a lot of work into making multiple streams seekable with shared timing info.) I mean bit-efficient, image-oriented, but image-encoding-neutral container formats.
There are at least two already-existing extensible image file formats that could be used for this: PNG and TIFF. In fact, TIFF was designed for this purpose — and even has several different encodings it supports!
But in practice, you don't see the people who create new image codecs these days thinking of themselves as creating image codecs — they think of themselves as creating vertically-integrated image formats-plus-codecs. You don't see the authors of these new image specifications thinking "maybe I should be neutral on container format for this codec, and instead just specify what the bitstream for the image data looks like and what metadata would need to be stored about said bitstream to decode it in the abstract; and leave containerizing it to someone else." Let alone do you ever see someone think "hey, maybe I should invent a codec... and then create multiple reference implementations for how it would be stored inside a TIFF container, a PNG container, an MKV container..."
It would be possible to define a JPEG XL payload for the HEIF container but it would not really bring anything except a few hundred bytes of extra header overhead and possibly some risk of patent infringement since the IP situation of HEIF is not super clear (Nokia claims it has relevant patents on it, and those are not expired yet).
Hey, thanks for the clarification! I was basing my info on Wikipedia (my bad), ISO BMFF page doesn't mention JXL at all, and even JPEG XL page has only small print in infobox saying that its "based on" ISO BMFF but the main article text doesn't mention that at all.
> But for JPEG XL this is not needed since JPEG XL already does have native support for all of these things — it was designed to be a still image codec after all
I suppose that is bit the thing grand-parent comment was complaining about, format not designed for general-purpose containers but rather as an standalone thing. I suppose it could be fun thought experiment to imagine what JXL would look like if it were specifically designed to be used in HEIF.
Of course it is well understandable that making tailored purpose-built format ends up better in many ways vs trying to fit into some existing generic thing.
> It would be possible to define a JPEG XL payload for the HEIF container but it would not really bring anything except a few hundred bytes of extra header overhead and possibly some risk of patent infringement since the IP situation of HEIF is not super clear (Nokia claims it has relevant patents on it, and those are not expired yet).
I suppose JXL-in-HEIF would allow some image management tools to have common code path for handling JXL and HEIC/AVIF files, grabbing metadata etc, and possibly would not need any specific JXL support. But that is probably not a practical concern in reality.
Videos has (most of the time at least) at least two tracks at the same time that has to be syncronized, and most of the time it's one video track and one audio track. With that in mind, it makes sense to wrap those in a "container" and allow the video and audio to be different formats. You also can have multiple audios/video tracks in one file, but I digress.
With images, it didn't make sense at least in the beginning, to have one container because you just have one image (or many, in the case of .gif).
OTOH a still image container would do nothing useful. If an image is all that needs to be contained, there's no need for a wrapper.
there are LOTS of uses for "image containers" that go beyond just pixels. heck, look at EXIF, which is extremely widespread - it's often stripped to save space on the web, but it's definitely useful and used.
It would, at least, create a codec-neutral location and format for image metadata, with codec-neutral (and ideally extensible + vendor-namespaced) fields. EXIF is just a JPEG thing. There is a reason that TIFF is still to this day used in medical imaging — it allows embedding of standardized medical-namespace metadata fields.
Also, presuming the container format itself is extensible, it would also allow the PNG approach to ancillary data embedding ("allow optional chunks with vendor-specific meanings, for data that can be useful to clients, but which image processors can know it's safe to strip without understanding because 'is optional' is a syntactic part of the chunk name") to be used with arbitrary images — in a way where those chunks can even survive the image being transcoded! (If you're unaware, when you transcode a video file between video codecs using e.g. Handbrake, ancillary data like thumbnail and subtitle tracks will be ported as-is to the new file, as long as the new container format also supports those tracks.)
Also, speaking of subtitle tracks, here's something most people may have never considered: you know how video containers can embed "soft" subtitle tracks? Why shouldn't images embed "soft" subtitle tracks, in multiple languages? Why shouldn't you expect your OS screen-reader feature to be able to read you your accessibility-enabled comic books in your native language — and in the right order (an order that, for comic books, a simple OCR-driven text extraction could never figure out)?
(There are community image-curation services that allow images to be user-annotated with soft subtitles; but they do it by storing the subtitle data outside of the image file, in a database; sending the subtitle data separately as an XHR response after the image-display view loads; and then overlaying the soft-subtitle interaction-regions onto the image using client-side Javascript. Which makes sense in a world where users are able to freely edit the subtitles... but in a world where the subtitles are burned into the image at publication time by the author or publisher, it should be the browser [or other image viewer] doing this overlaying! Saving the image file should save the soft subtitles along with it! Just like when right-click-Save-ing a <video> element!)
That would be a layered image format, like .psd (Photoshop).
It's an interesting idea, memes could become editable :)
In the alt-text case specifically, you could allow for optional styling info so that the gloss can be laid out as a visual replacement for the original text that was on the page. But that's not really necessary, and might even be counterproductive to some use-cases (like when interpretation of the meaning of the text depends on details of typography/calligraphy that can't be conveyed by the gloss, and so the user needs to see the original text with the gloss side-by-side; or when the gloss is a translation and the original is written with poetic meter, such that the user wants the gloss for understanding the words but the original for appreciating the poesy of the work.)
Concrete use-cases:
• the "cleaner" and "layout" roles in the (digitally-distributed) manga localization process, only continue to exist, because soft-subbed images (as standalone documents) aren't a thing. Nobody who has any respect for art wants to be "destructively restoring" an artist's original work and vision, just to translate some text within that work. They'd much rather be able to just hand you the original work, untouched, with some translation "sticky notes" on top, that you can toggle on and off.
• in the case of webcomic images that have a textual "bonus joke" (e.g. XKCD, Dinosaur Comics), where this is currently implemented as alt/title-attribute text — this could be moved into the image itself as a whole-image annotation, such that the "bonus joke" would be archivally preserved alongside the image document.
“The Plain Text Extension contains textual data and the parameters necessary to render that data as a graphic, in a simple form. The textual data will be encoded with the 7-bit printable ASCII characters. Text data are rendered using a grid of character cells defined by the parameters in the block fields. Each character is rendered in an individual cell. The textual data in this block is to be rendered as mono-spaced characters, one character per cell, with a best fitting font and size.”
“The Comment Extension contains textual information which is not part of the actual graphics in the GIF Data Stream. It is suitable for including comments about the graphics, credits, descriptions or any other type of non-control and non-graphic data.”
I hesitate to say GIF89a "supported" it since in practice approximately zero percent of software can use either extension. `gIFt` was dropped from the PNG spec for this reason: https://w3c.github.io/PNG-spec/extensions/Overview.html#DC.g...
If it had been well-supported we might have avoided the whole GIF pronunciation war. Load up http://cd.textfiles.com/arcadebbs/GIFS/BOB-89A.GIF in http://ata4.github.io/gifiddle/ and check out the last frame :)
- contain multiple streams of synced video, audio, and subtitles
- contain alternate streams of audio
- contain chapter information
- contain metadata such as artist information
For web distribution of static images, you want almost none of those things, especially regarding alternate streams. You just want to download the one stream you want. Easiest way to do that is to just serve each stream as a separate file, and not mux different streams into a single container in the first place.
Also, I could be wrong on this part, but my understanding is that for web streaming video, you don't really want those mkv* features either. You typically serve individual and separate streams of video, audio, and text, sourced from separate files, and your player/browser syncs them. The alternative would be unnecessary demux on the server side, or the client unnecessarily downloads irrelevant streams.
The metadata is the only case where I see the potential benefit of a single container format.
* Not specific to mkv, other containers have them of course
That's one school of thought. Some of the biggest streaming providers simply serve a single muxed video+audio HLS stream based on bandwidth detection. Doesn't work very well for multi-language prerecorded content of course, but that's just one use case.
My understanding is that YouTube supports both the "separate streams" and "specific mux per-bandwidth profile" methods, and picks one based on the codec support/preferences of the client.
HTTP file transfer protocols support partial downloads. A client can choose just to not to download irrelevant audio. I think most common web platforms already work this way, when you open a video it is likely to be in .mp4 format, and you need to get the end of it to play it, so your browser gets that part first. I am not entirely sure.
This is not a good assumption. MKV supports a loooot of things which many video players will not support at all.
And IIRC some browsers do not support MKV.
So you’re absolutely going to see TIFF containers with JPEG or JPEG2000 tiles used for geospatial, medical, or hi-res scanned images, but given the sad state of open tooling for all of these, there’s little to no compatibility between their various subsets of the TIFF spec, especially across vendors, and more or less no FOSS beyond libtiff. (Not even viewers for larger-than-RAM images!) Some other people have used TIFF but in places where’s very little to be gained from compatibility (e.g. Canon’s CR2 raw images are TIFF-based, but nobody cares). LogLuv TIFF is a viable HDR format, but it’s in an awkward place between the hobby-renderer-friendly Radiance HDR, the Pixar-backed OpenEXR, and whatever consumer photo thing each of the major vendors is pushing this month; it also doesn’t have a bit-level spec so much as a couple of journal articles and some code in libtiff.
Why did this happen? Aside from the niche character of very large images, Adobe has abandoned the TIFF spec fairly quickly after it acquired it as part of Aldus, but IIUC for the first decade or so of that neglect Adobe legal was nevertheless fairly proactive about shutting up anyone who used the trademarked name for an incompatible extension (like TIFF64—and nowadays if you need TIFF you likely have >2G of data). Admittedly TIFF is also an overly flexible mess, but then so are Matroska (thus the need for the WebM profile of it) and QuickTime/BMFF (thus 3GPP, MOV, MP4, ..., which are vaguely speaking all subsets of the same thing).
One way or another, TIFF is to some extent what you want, but it doesn’t get a lot of use these days. No browser support either, which is likely important. Maybe the HEIF container (yet another QuickTime/BMFF profile) is better from a technical standpoint, but the transitive closure of the relevant ISO specs likely comes at $10k or more. So it’s a bit sad all around.
Webm went way too far when they stripped out support for subtitles. The engineers who made that decision should be ashamed.
[1] https://www.webmproject.org/docs/container/#webvtt-guideline...
[2] https://trac.ffmpeg.org/ticket/5641
[3] https://matroska.org/technical/codec_specs.html#s_textwebvtt
Fun fact: several broadcast standards use Bitstream TrueDoc Portable Font Resource, which was supported for embedded web fonts way back in Netscape 4:
https://people.apache.org/~jim/NewArchitect/webrevu/1997/11_...
https://web.archive.org/web/20040407162455/http://www.bitstr...
“The PFR specification defines the Bitstream portable font resource (PFR), which is a compact, platform-independent format for representing high-quality, scalable outline fonts.
Many independent organizations responsible for setting digital TV standards have adopted the PFR font format as their standard font format, including:
— ATSC (Advanced Television Systems Committee)
— DAVIC (Digital Audio Visual Council)
— DVB (Digital Video Broadcasting)
— DTG (Digital TV Group)
— MHP (Multimedia Home Platform)
— ISO/IEC 16500-6:1999
— OCAP (OpenCable Application Platform)”
[1] https://cve.mitre.org/cgi-bin/cvekey.cgi?keyword=libtiff
[2] https://exiftool.org/#supported (search for "TIFF-based")
That’s an impressive number of CVEs for a fairly modest piece of code, although the sheer number of them dated ≥ 2022 baffles me—has a high-profile target started using libtiff recently, or has some hero set up a fuzzer? In any case libtiff is surprisingly nice to use but very old and not that carefully coded, so I’m not shocked.
I’m not sure about the absolute offsets, though. In which respect are those more error-prone? If I was coding a TIFF library in C against ISO or POSIX APIs—and without overflow-detecting arithmetic from GCC or C23—I’d probably prefer to deal with absolute offsets rather than relative ones, just to avoid an extra potentially-overflowing addition whenever I needed an absolute offset for some reason.
There are things I dislike about TIFF, including security-relevant ones. (Perhaps, for example, it’d be better to use a sequential format with some offsets on top, and not TIFF’s sea of offsets with hopefully some sequencing to them. Possibly ISO BMFF is in fact better here; I wouldn’t know, because—well—ISO.) But I don’t understand this particular charge.
I think parsing file format with absolute offsets is similar to handling a programming language with all GOTOs, compared to relative offsets which are more like structured control flow.
1. put a reference to the decoder into the header of the compressed file
2. download the decoder only when needed, and cache it if required
3. run the decoder inside a sandbox
4. allow multiple implementations, based on hardware, but at least one reference implementation that runs everywhere
Then we never need any new formats. The system will just support any format. When you're not online, make sure you cached the decoders for whatever files you installed on your system.
Apart from that of course the decoder has to be fast and thus native and interface with the OS, so the decoder is X86 on the today version of Windows, until the company hosting it dies and the patented, copyrighted decoder disappears from the internet.
A decoder can be extremely isolated. It's much easier to sandbox a decoder than to sandbox javascript, for example.
If we had standard headers, then reading metadata wouldn't be part of the decoder. The decoder would only need to take in bytes and output a bitmap, or take in bytes and output PCM audio. It doesn't need to be able to call any functions, or run any system calls, and the data it outputs can safely contain any bytes because nothing will interpret it.
It's like taking the very core of webassembly and then not attaching it to anything. The attack surface is astoundingly small.
You just need to give it a big array of memory and let it run arithmetic within that array, plus some control flow instructions. Easy to interpret, easy to safely JIT compile.
The part of sandboxing that's hard is dealing with I/O, or giving useful tools to the sandboxed code, or implementing data structures for the sandboxed code. You don't need any of that for a multimedia decoder. You just let it manipulate its big block of bytes, and make sure you bounds check.
A Java VM exposes tens of thousands of functions to the code inside it. A barebones sandbox exposes zero. It just waits for the HLT opcode.
And when it gives you raw RGB data, or raw PCM data, there's no way to hide a triggerable malicious payload inside. If the code does something bad, the worst it can do is show you the wrong image.
But my suggestion would be that you build a video codec out of it. Preferably one that has the properties the market demands: performance and energy efficiency.
You'd use existing codecs, and the way you get good performance and energy efficiency on a video codec is by having a hardware implementation. Software decoding doesn't even come into that picture.
As far as practical software decoding outside of battery-powered video, can I just point at webassembly? Especially the upcoming version with vector instructions. You could use normal webassembly, or even an extra-restricted version. It gets pretty good performance, and when you remove its ability to talk to the outside world it goes from pretty good to extremely good security.
WebAssembly codecs indeed exist, and they are impractical due to a lack in performance.
The container parser would not be dynamically downloaded, and may or may not be sandboxed.
We don't need a new container with almost every codec. We just need the new codec itself.
> WebAssembly codecs indeed exist, and they are impractical due to a lack in performance.
Mostly because they don't have vector instructions yet, I bet. But plenty of webassembly is within 50% of native, which is good enough for lots of things, which includes image decoding for sure.
The challenge remains for you to actually provide the codec you describe. Which a few comments ago was trivial because it was a hardware codec anyway, now it’s just a bit of WebAssembly away. Well that should be trivial because cross compilers to WebAssembly exist. So why don’t you just provide a few real world examples? Your probably not the first to think of these ideas, there has to be a reason why it hasn’t been done yet.
Not "magically". But you only need one or two, and they don't need to be very fast, so you can put a lot of effort into making them secure.
But more importantly, browsers already have many container decoders. This is not an expansion in attack surface. The goal here is allowing a lot more codecs compared to current browsers without a significant increase in attack surface compared to current browsers. Pointing out flaws that already exist doesn't disqualify the idea.
> So why don’t you just provide a few real world examples? Your probably not the first to think of these ideas, there has to be a reason why it hasn’t been done yet.
Image decoders in webassembly already exist. Did you even look? Including JXL!
Video decoding needs more support structure in the browser, and I already said some decoders need things that are being added to webassembly but aren't done yet. Even then, the first google result for "av1 webassembly" is a working decoder from five years ago.
You no longer need "printer drivers", they're supposed to be automatically downloaded, installed and ran in a sandbox. You never need any "new drivers". The system will support any printer.
Except the "sandbox" was pretty weak and full of holes.
Nothing prevents you from installing only the trusted ones.
Second, software is getting so complicated that if we don't build secure sandboxes anyway then at some point people will be bitten by a supply chain attack.
The ISOBMFF format is used as a container for MP4, JPEG 2000, JPEG XL, HEIF, AVIF, etc.
And yes, there are ffmpeg-like "uber-libraries" for images: ImageMagick, GraphicsMagic, libvips, imlib2 and gdk-pixbuf are examples of those. They support basically all image formats, and applications based on one of these will 'automatically' get JPEG XL support.
Apple also has such an "uber-library" called CoreMedia, which means any application that uses this library will also get JPEG XL support automatically.
you can also with a simple mod on a older commit of ffmpeg (the commit that added animated jxl broke this method and I haven't gotten around to fixing it) by simply adding jxl 4cc to mux JXL sequences into MKV.
We are still using image formats from the 90s, and their matching containers, and they are good enough, so there is not much work for going beyond that. There is no real incentive for making a more flexible format. By comparison, video is the biggest bandwidth hog and people care a lot.
And mkv supports video, multiple sound tracks, subtitles,... All using different codecs made by different people (ex: h265+opus or vp9+vorbis, or any other combination). An image container usually only has the image and a few metadata.
It loads JXL if your client supports it.
Recent builds of Chrome and Edge now support and display JXL on iOS 17. They have to use the Safari engine underneath, but previously they suppressed JXL, or maybe the shared engine did.
See 2.5.6 here - https://developer.apple.com/app-store/review/guidelines/
And not only that, it's reasonably fast to encode on consumer hardware.
JPEG XL has the ambition to supplant all the images formats of the next 20+years.
You will have to define what is marginally better. WebP is definitely marginally better than JPEG. And JPEG XL is easily 30- 40% BD-Rate at the same quality at BPP 0.8 or over.
Second Life went with JPEG2000 for textures, and when they open sourced the client, they had to switch to an open source library that was dog slow. Going into a new area pretty much froze the client for several minutes until the textures finally got decoded.
That’s not true. JPEG 200 had substantially smaller file sizes and better progressive decoding – something like responsive images could have just been an attribute telling your browser how many bytes to request for a given resolution. It also had numerous technical benefits for certain types of images - one codec could handle bitonal images more efficiently than GIF, lossless compressed better than TIFF, effortlessly handle colorspaces and bit depths we’re just starting to use on the web, etc.
What doomed it was clumsy attempts to extract as much license revenue as possible. The companies behind it assumed adoption was inevitable so everything was expensive - pay thousands for the spec, commercial codecs charged pretty high rates, etc. and everyone was so busy pfaffing around with that that they forgot to work on things like interoperability or performance until the 2010s. Faced with paying money to deal with that, most people didn’t and the market moved on with only a few exceptions like certain medical imaging or archival image applications. In the early 2000s the cost of storage and disk/network bandwidth meant you could maybe try to see the numbers as plausibly break-even but over time that faded while the hassle of dealing with the format did not.
You get better compression and services can deliver multiple resolutions/qualities from the same stored image (reducing storage or compute costs), all transparent to the user.
So your average user will not care but your cloud and web service companies will. They are going to want to adopt this tech once there's widespread support so they can reduce operating costs.
Amazon alone sold over 375 million items this last Prime Day. Let’s say that was 200 million items loaded/day (ignoring the unpublished number of failed sales), with 9 images (the maximum from a cursory glance) at 2000x2000 for a 1:1 ratio and zoomability. For a 90% quality JPG at 24-bit color that’s 410KB. ((410KBx9)200,000,000) = 738TB. Now imagine cutting that in half with no perceptive difference except faster loading to the end-user.
For end users the other options may be more desirable, but I would argue the importance is in the compression itself.
(Having said that I do wish for JPEG XL to become a true successor)
so do we think Chrome will reverse their decision to drop support?
Maybe with a banner like "You are using Chrome. You might have a degraded experience due to the lack of support for better image formats. Consider Firefox".
In either case: Have Chrome telemetry report home that "user could have 20% faster page load with JXL support".
Nothing changes for Chrome users, especially sites using the <picture> element where the first supported image format is used.
I’m sure YouTubers and tech sites will love to do Safari vs Chrome (and Co.) content to spread the message that Chrome is inferior.
(WASM) Polyfill and we're done.
Any resemblance with previous events would be totally unintentional, of course :-)
Nope.
Microsoft could probably push Google over the Edge. They have a lot of influence over Chrome with Edge/Windows defaults, business apps and such.
(and no I can’t tell you how I know this)
It feels disturbingly tribal.
The argument was that there's no industry support (apparently this means: beyond words in an issue tracker), let's see how acceptance is with Safari supporting it.
An uptick in JXL use sounds like a good-enough reason to re-add JXL support, this time not behind an experimental flag. Maybe Firefox even decides to provide it without a flag and in their regular user build.
There are arguments for the new format, but the Chrome people seemed unwilling to maintain support for it when pick-up was non-existent (Firefox could have moved it out of their purgatory. Safari could have implemented it earlier. Edge could have enabled it by default. Sites could use polyfills to demonstrate that they want the desirable properties. And so on.)
To me, the situation was one of "If Chrome enables it, people will whine how Chrome forces file formats onto everybody, making the web platform harder to reimplement, a clear signal of domination. If they don't enable it, people will whine how Chrome doesn't push the format, a clear signal of domination", and they chose to use the variant of the lose-lose scenario that means less work down the road.
Of course there is no pick-up when Chrome, with its massive market share, doesn't support it. Demanding pick-up before support makes no sense for an entity with such a large dominance.
- Microsoft enabling the flag in Edge by default and telling people that websites can be 30% smaller/faster in Edge, automatically adding JXL conversion in their web frameworks
- Apple doing the same with Safari (what they're _now_ doing)
- Mozilla doing the same with Firefox (instead of hiding that feature in a developer-only build behind a flag)
None of that happened so far, only the mixed signal of "lead and we'll follow" and "you are too powerful, stop dominating us." in some issue tracker _after_ the code has been removed.
> "you are too powerful, stop dominating us."
That's twisting things. The problem was that the argument of the Chrome team against JPEG XL was self-refuting. They were themselves the main cause of what they complained about.
Chrome had that code, hidden behind a flag. There wasn't any kind of activity. No questions "when will you put it in by default in Chrome?". No other Blink-based browser (Edge, Brave, Vivaldi, Opera) that could easily pick up the support by enabling that damn flag by default did so. Firefox hid JXL support even better than Chrome. No image sharing site that did the math and considered "200KB for a polyfill saves us and our users megabytes in traffic on each visit" and acted on that.
That doesn't look like anybody is interested in JXL support.
I'm bringing this up again and again because I dislike that notion of "Chrome is the market leader and we're powerless to do anything about it. Bad Google." It neither encourages the Chrome folks to do better nor anybody else to pick up the slack. It's 100% complaint, no matter what Chrome does.
AVIF works extremely well at compressing images down to very small sizes with minimal losses in quality but loses comparatively to JPEG XL when it comes to compression at higher quality. Also I believe AVIF has an upper limit on canvas sizes (2^16 pixels by 2^16 pixels I think) where JEPGXL doesn't have that limitation.
Also existing JPEGs can be losslessly migrated to JPEGXL which is preferable to a lossy conversion to AVIF.
So it's preferable to have JPEG XL, webP, and AVIF.
- webP fills the PNG role while providing better lossless compression
- AVIF fills the JPEG role for most of your standard web content.
- JPEG XL migrates old JPEG content to get most of the benefits of JPEG XL or AVIF without lossy conversion.
- JPEG XL fills your very-high fidelity image role (currently filled by very large JPEGs or uncompressed TIFFs) while providing very good lossless and lossy compression options.
If you have image-heavy workflows and care about storage and/or bandwidth then JPEG-XL pairs great with AVIF: JPEG-XL is great for originals and detail views due to its great performance at high quality settings and high resolution support, meanwhile AVIF excels at thumbnails where resolution doesn't matter and you need good performance at low quality settings.
I'm hoping it gets adopted as a better underlying technology for various RAW formats, and hopefully a better successor to the DNG format while we're at it (currently these are TIFF based). I'm not even a professional photographer, and my hard drive is still mostly occupied by RAW files.
Specifically, it matters for source files and intermediate files.
With RAW files from the camera, the higher the bit depth of the analog-to-digital conversion (ADC) step, the less posterization this introduces on the signal. Theoretically at least, you're still limited by the sensor's dynamic range, and there are other subtleties involved, like light perception being logarithmic instead of linear, but RAW encodings being linear[0][1]. But in simple terms: paired with a sensor with high dynamic range and good ADC, a higher bit depth results in less noise and higher dynamic range. Which allows one to recover more fine detail from shadows and highlights. Which makes the camera more forgiving in normally difficult lighting scenes (low light and/or high contrast). So a higher bit depth can aid in giving photographers creative freedom when shooting, and more flexibility in editing their photos without loss of fidelity.
So yes, it is an important cog in the machine that is the whole processing pipeline.
Having said that, as I mentioned our eyes perceive light logarithmically. The dynamic range of the human eye is... complicated to determine, because it adjusts so quickly. At night it may go up to 20 stops, during the day 14 stops is likely to be the typical range[2]. So it's probably not a coincidence that digital cameras have "stalled" at using 14 bits for their RAW files, typically: the photographer likely wouldn't be able to see more contrast in the lights and shadows before taking a photo anyway!
[0] https://www.dpreview.com/articles/4653441881/bit-depth-is-ab...
[1] No I don't understand why floating point ADCs aren't used either, seems like it would be a more sensible approach to me and they do exist: https://ieeexplore.ieee.org/abstract/document/776106
That's tremendously simpler, both from an architectural and maintenance standpoint (for any site that deals with images), than what you would usually have to do, such as relying on either a third party host (and added cost, latency (without caching), and potential downtime/outage) or pushing it through the (very terrible and memory/cpu-wasteful codebase at this point) ImageMagick/GraphicsMagick library (and potentially managing that conversion as a background job which incurs additional maintenance overhead), or getting VIPS to actually successfully build in your CI/CD workflow (an issue I struggled with in the past while trying to get away from "ImageTragick").
You get to chuck ALL of that and simply hold onto the originals in your choice of stateful store (S3, DB, etc.), possibly caching it locally to the webserver, and just... compute the number of pixels you need given the requested dimensions (which is basically just: ((requested x)*(requested y))/((full-size x)*(full-size y)) percentage of the total binary size, capping at 100%), and bam, truncate.
Having built out multiple image-scaling (and caching, and sometimes third-party-hosted) workflows at this point, this is a very attractive feature, speaking as a developer.
The unique part AFAIK is that you can order the data blocks however you want, allowing progressive loading that prioritizes more important or higher detailed areas: https://opensource.googleblog.com/2021/09/using-saliency-in-...
(your thumbnails may or may not look terrible this way, as well. really better suited for progressive loading)
https://www.youtube.com/watch?v=UphN1_7nP8U
This comparison video is admittedly a little unfair though, because AVIF would have easily 30% lower file size than JPEG XL on ordinary images with medium quality.
- JPEG XL can do lossless compression better than PNG if I’m right.
- At low bit rates, JPEG XL isn’t that far from AVIF quality. You will only use it for less important stuff like “decorations” and previews anyway so we can be less picky about the quality.
- For the main content, you will want high bit rates which is where JPEG XL excels.
- Legacy JPEG can be converted to JPEG XL for space savings at no quality loss.
The use cases of WebP is limited, the actual advantage over decent JPEG and isn't that big, and unless you use a lot of lossless PNG I would argue it should have never been pushed as the replacement of JPEG. To this day I still dont know why people are happy about WebP.
According to Google Chrome, 80% of images transferred has an BPP 1.0 or above. The so called "low bit rate" happens at below BPP 0.5. The current JPEG XL is still no optimised for low bitrate. And judging from the author's tweet I dont think they intend to do it any time soon. And I can understand why.
"Can't decide between a Nissan March, Jpeg SX, or the Honda Jazz EX"
It might be I am not encoding properly but when I did trials with a small number of photos with the goal of compressing pictures I took with my Sony α7ii at high quality I came to the conclusion that WEBP was consistently better than JPEG but AVIF was not better than WEBP. I did think AVIF came out ahead at lower qualities as you might use for a hero image for a blog.
Lately I've been thinking about publishing wide color gamut images to the web, this started out with my discovery that a (roughly) Adobe RGB monitor adds red when you ask for an sRGB green because the sRGB green is yellower than the Adobe RGB green and this is disasterous if you are making red-cyan stereograms.
Once I got this phenomenon under control I got interested in publishing my flat photos in wide color gamut, I usually process in ProPhotoRGB so the first part is straightforward. A lot of mobile devices are close to Display P3, many TV sets and newer monitors approach Rec 2020 but I don't think cover it that well except for a crazy expensive monitor from Dolby.
Color space diagram here: https://en.wikipedia.org/wiki/Rec._2020#/media/File:CIE1931x...
Adobe RGB and Display P3 aren't much bigger than the sRGB space so they still work OK with 8-bit color channels but if you want to work in ProPhotoRGB or Rec 2020 you really need more bits, my mastering is done in 16 bits but to publish people usually use 10-bit or 12-bit formats which has re-awakened my interest in AVIF and JPEG XL.
I'm not so sure if it is worth it though because the space of colors that appear in natural scenes is a only bit bigger than sRGB
https://tftcentral.co.uk/articles/pointers_gamut
but much smaller than space of colors that you could perceive in theory (like the green of a green laser pointer. Definitely Adobe RGB covers the colors you can print with a CMYK process well, but people aren't screaming out for extreme colors although I expect to increasingly be able to deliver them. So on one hand I am thinking of how to use those colors in a meaningful way but also the risk of screwing up my images with glitchy software.
In practice that larger space of things you could perceive "in theory" is full of everyday phenomena, and very brilliant colors and HDR scenes (e.g. fireworks against a dark sky) tend to be something people particularly enjoy looking at/taking pictures of.
This surprises me greatly if you're talking about image quality. I've always found WebP to be consistently worse than JPEG in quality.
I only use WebP for lossless images, because at least then being smaller than PNG is an advantage.
A lot of monitors from most vendors support Display P3, even if it is usually named slightly erroneously as DCI P3.
Display P3 differs from the original DCI P3 specification by having the same white color and the same gamma as sRGB, which is convenient for the manufacturers because all such monitors can be switched between the sRGB mode (which is normally the default mode) and the Display P3 mode.
Nonetheless, even if today most people that have something better than junk sRGB displays have Display P3 monitors (many even without knowing this, because they have not attempted to change the default sRGB color space of their monitors), images or movies should be distributed as you say, using the Rec. 2020 color space, so that those with the best displays shall be able to see the best available quality of the image, while the others will be able to see an image with a quality as good as allowed by their displays.
Adobe RGB was conceived for printing better images and it is not useful on monitors because it does not correct the main defect of sRGB, which is the red.
Moreover, if I switch my Dell Display P3 monitor (U2720Q) from 30-bit color to 24-bit color, it becomes obviously worse.
So, at least in my experience, 10-bit per color component is always necessary for Display P3 in order to benefit from its improvements, and on monitors there is a very visible difference between Display P3 (or DCI P3) and sRGB.
There are a lot of red objects that you can see every day and which have a more saturated red than what can be reproduced by an sRGB monitor, e.g. clothes, flowers or even blood.
For distributing images or movies, I agree that the Rec. 2020 color space is the right choice, even if only few people have laser projectors that can reproduce the entire Rec. 2020 color space.
The few with appropriate devices can reproduce the images as distributed, while for the others it is very simple to convert the color space, unlike in the case when the images are distributed in an obsolete color space like sRGB, or even Adobe RGB, when all those with better displays are still forced to view an image with inferior quality.
They adopted HEIF, and have not adopted AV1 video.
(At least Safari 16.5.2 on Ventura 13.4.1 won't open .heif / .heic files for me).
Sure, Apple shipped the first consumer computer that supported Display P3 in 2015 [1].
And while there are several other vendors including Google with devices that support Display P3, Apple’s 2 billion devices is not nothin’.
Note that jpg xl is different from jpg 2000 and jpg xr
Well, at least in the tiny part of the IT world I get to control, I always try to validate based on both the three letter extension and any common or sensible expansion of that. So ".jpg" or ".jpeg", ".jxl" or ".jpegxl" etc. etc. (And in most cases, I actually try to parse the binary itself, because you can't trust the extension much anyway.)
"Well, three characters is the bare minimum. If you feel that three characters is enough, then okay. But some people choose to have longer filename extensions, and we encourage that, okay? You do want to express yourself, don't you?"
I mean, it feels like the same static image codec should be used in whatever free standard is being pushed for both video I-frames and images, since the problem is basically the same.
So crazy to me that Apple and Google fight over image formats like this.
I guess this is just the next round.
"""Google PIK + Cloudinary FUIF = JPEG XL"""
Before saying what it is, was a little of-putting.
To understand that article required me searching for FUIF [1], PIK [2] and a brief explanation of what JPEG XL is trying to achieve.
I double down on my "complaint" - I'd call it constructive criticism - that article was poorly written. It's actually quite a good story that their Free Universal Image Format (FUIF) has achieved what it has. That's a great acronym, especially for a world that thinks JPEG XL is a good acronym! Why not put in in the article.
To save anyone else time:
[1] https://github.com/cloudinary/fuif [2] https://github.com/google/pik
However if you manually request that image with a custom `image/jxl` at the start of the `Accept:` header, you get a JPEG XL result. So GP is correct, but you won't see that behavior except on their PC (errr, Mac) -- unless you use Firefox and enable JPEGXL support in about:config, of course.
image/webp,image/avif,image/jxl,image/heic,image/heic-sequence,video/;q=0.8,image/png,image/svg+xml,image/;q=0.8,/;q=0.5
https://github.com/mozilla/mozjpeg#mozilla-jpeg-encoder-proj...
Initially, a new software codec will grind the cpu and battery-life like its on a 20 year old phone. Then often becomes pipelined into premium GPUs for fringe users, and finally mainstreamed by mobile publishers to save quality/bandwidth when the market is viable (i.e. above 80% of users).
If anyone thinks they can shortcut this process, or repeat a lock-down of the market with 1990s licensing models... than it will end badly for the project. There are decades of media content and free codecs keeping the distribution standards firmly anchored in compatibility mode. These popular choices become entrenched as old Patents expire on "good-enough" popular formats.
Best of luck, =)
I don't think we have much deployment of WebP or AVIF hardware decoders yet the formats have widespread use and adoption.
What I am seeing is the high-water-mark for WebP was 3 years ago... It is already dead as a format, but may find a niche use-case for some users.
Consider built-in web-cams that have hardware h264 codecs built-in, as the main cpu just has to stream the data with 1 core. Better battery life for both the sender and receiver.
Keep in mind a single web page may have hundreds of images, and mjpeg streams are still popular in machine-vision use-cases. As most media/gpu hardware is now integrated into most modern browsers, the inertia of "good-enough" will likely remain.. and become permanent as patents expire. =)
The choice a chip maker has is to include popular legacy/free codecs like mp3, or pay some IP holder that won't even pick up a phone unless there is over $3m on the table. h264 was easy by comparison, but hardly ideal. NVIDIA was a miracle when you consider what they likely had to endure. =)
Is hardware avif decoding done anywhere? The only example I can think of where this is done is HEIF on iOS devices, maybe.
Some cloud GPUs have jpeg decoding blocks for ingesting tons of images, but that'd not really the same thing.
mjpeg streams are still popular in machine-vision use-cases. And you can be fairly certain your webcam has that codec built-in if the device was made in the past 15 years. Note most media/gpu hardware is now integrated into modern browsers. =)