John Carmack on JPEG
twitter.com
twitter.com
JPEG 2000 is not used much. Decoding is slow, and encoding is slower. The decoders are either buggy or proprietary. It has way too many options internally. The big users of JPEG 2000 are medical. It has a "lossless" mode, and medical imagery is usually stored lossless because you really don't want compression artifacts in X-rays.
(I've been struggling with JPEG 2000 recently. Second Life and Open Simulator use it for asset storage. Some images won't decompress properly with OpenJPEG, a free decoder. I finally found out why. There's a field used for "personal health information" in JPEG 2000. This is where the patient name and such go in a CAT scan. That feature was added after the original version. Some older images apparently have junk in that field, which causes problems.)
You really don't want compression artifacts in any image, which is why you should set the compression ratio wisely when encoding. I don't see how X-rays are any different.
https://www.theregister.com/2013/08/06/xerox_copier_flaw_mea...
JBIG2 is a format for storing black and white documents in a highly compressed way. It works by detecting each letter in the document, and then replacing it with a pointer to the reference version of that letter, up to a certain threshold. Basically compression via OCR.
Of course, this means that when a distorted letter is too close to the reference version of another letter, it will get replaced with a clean version of that incorrect one. So even though a human could easily recognize that something was off with that letter in the original image, the JBIG2-compressed image has no such clue!
What’s really bad is that JBIG2 compression was built into certain Xerox machines that were used by archivists to digitize important documents for years until someone noticed the discrepancies. JBIG2 was promptly banned for archival purposes, but there might still be a ton of documents with these kind of invisible errors in our archives! :-)
On the other hand, JBIG2 doesn’t actually do OCR. It only does template matching of similar-looking blocks of pixels. The compressor doesn’t try to understand which letter those pixels represent.
There's one big win though. Digital cinema projectors, in now pretty much all theatres, run DCPs which have movies encoded with it.
Are you aware of the downsides?
https://www.dkriesel.com/en/blog/2013/0802_xerox-workcentres...
It stands to reason that not only Xerox had problems with JBIG2.
JPEG2000 is such a cool choice for Second Life. The fact any truncation of a JPEG2000 bitstream is just a lower-resolution version of the image makes for convenient progressive enhancement when loading textures over the network, right?
(And Second Life even used JPEG2000 for geometry, sort-of. I guess with the advent of proper mesh support, that may be less common now though.)
Dedicated texture compression formats would also do this since they use mipmaps, but I don't know if you can stream those.
What do you mean by this? Zip, 7zip, gzip, zstd, etc. get little to no compression on PNG. JPEG, or JPEG2000. Presumably you mean something other than zipping up directories of images.
I work in the library/archiving space where people spent years trying to make this happen due to the compression wins, support for a good range of color depths and spaces, and the progressive decoding is perfect for browsing galleries of high-res images before zooming way into the one you wanted.
The frictional cost largely canceled that out: people don’t trust unreliable formats and the JP2 files which only opened in one of {Kakadu, Aware, Adobe} and couldn’t be used with any open source tools live long in the memory. Performance in Jasper was wretched … and while the vendors thought that’d boost sales, it seemed far more effective to me at getting people to use other formats instead.
These days, if anyone is talking about new image formats the first question is what they’re doing for open source (especially things like ImageMagick which half the world uses) and specifically browsers. A WASM polyfill for <picture> or <img srcset> is really critical because it means people don’t have to transcode everything for a handful of users.
Original image: https://www.fnordware.com/j2k/jp2samples.html
What it looks like in GIMP: https://imgur.com/a/HCNz7ga
1. Per Photoshop, "The embedded ICC profile cannot be used because the ICC profile is invalid."
2. After using the OS X Preview.app "Assign Profile" command to replace the embedded ICC profile with the system-supplied sRGB profile, the image displays correctly in Safari.
3. The ImageMagick display command is apparently ignoring the embedded profile and assuming sRGB, as it continues to display the image correctly even when I replace the embedded profile with an obviously incorrect (but valid and correctly interpreted by Photoshop and Safari) profile.
What possesed someone to actually attempt to use JPEG2000 I do not understand.
JPEG has introduced multiple new formats since JPEG 2000, all of which are better, such as JPEG XR and JPEG XL.
Apple has decided it shouldn't be by not allowing its use on iOS.
There are various problems in XVideo that make it hard to use for general application display, but the fact that this was a thing they included in it makes me think that a substantial number of video cards have included YUV decoding capability in hardware since the previous millennium.
Namely, VK_KHR_sampler_ycbcr_conversion. I'm actually not sure if other APIs provide the equivalent yet, the HW capability is pretty recent...
GL_EXT_YUV_target is close, but it doesn't transparently include the specific step of YUV -> RGB conversion, just the sampling multiple planes + upsampling for chroma (and convenience functions for shader conversion.) (I think? actually not certain, since I can't find definition of how chroma is upsampled)
Well, I guess Vulkan's isn't completely transparent, since the conversion sampler has additional restrictions as to how it can be used. But at least with it shaders can be written to not care whether the source is RGB or not.
[1] https://developer.android.com/reference/android/graphics/Col...
while not standard, writing a shader to decode is not hard(admittedly if you know shaders)
The time I spent dealing with video decode/encode was not pleasurable but interesting.
The really hard part is the actual encoding/decoding algorhtims
So it's analogous, but there's a critically important difference. (Not that you claimed otherwise.)
Ah, the old "Step 2: Draw the rest of the owl" scenario.
https://wiki.mamedev.org/index.php/Driver:Apple_II (which had no RGB at all)
https://old.reddit.com/r/OpenEmu/comments/hhxnry/sharing_my_... (showing Sonic, among others, though I'm not sure if this is really about NTSC artifacting or just dithering)
http://forum.arcadecontrols.com/index.php/topic,152465.msg15...
https://emulation.gametechwiki.com/index.php/NTSC_Filters
So, there are some exceptions, including a few extremely popular platforms, but certainly you are correct that emulating the majority of game platforms does not need NTSC artifacting.
Carmack is giving an extremely simplified version of the world as it exists today. Your hardware can already use YUV formats, maybe it is even outputting YUV because you chose a resolution and bit depth that forced it to for bandwidth or clock constraints. Complex apps like browsers that do their own compositing can already choose to use YCbCr formats when supported for textures.
You are correct, and you can find details on that by reading the card documentation (as far as it is available). That was in particular the case when cards had a fixed-function rendering pipeline.
With the advent of more general purpose cores on the GPU for the rendering pipeline, YUV decoding has been offloaded to them.
He mentions 4:2:0 chroma subsampling. But he doesn't mention chroma siting. Or alternative sumbsampling schemes. Or matrix coefficients. Or full-range vs video-range (a.k.a. JPEG vs MPEG range). Heck, how you even arrange subsampled data varies by system (many libraries like planar; Apple likes bi-planar; etc.).
I'd love to see more support for rendering subsampled Y'CbCr formats so you don't have to use so much RAM, but it gets complicated quick.
A lot of Y'CbCr -> RGB converters actually disagree with each other. They're all close enough that casual users don't notice or care about the small discrepancies.
Again about what to do when missing or conflicted metadata is available. :)
The point of my post was to help casual readers know that this gets complicated fast. Someone might think defining FMT_JPEG_YUV is easier and simpler than it actually is. I don't fault John for that. It's not a fault of anyone, really.
So since you already have to process the image, it doesn't seem like a big ask to convert from JPEG-flavored YUV to GPU-flavored YUV. But I'm not an expert, so maybe this is hard/lossy?
It's not lossy. But Y'CbCr -> RGB is not as simple as you might think. My whole post was about Y'CbCr -> RGB conversion. It's doable; I'm not saying it's impossible. There's just several various flavors of Y'CbCr and correctly handling all of them (or at least the majority/most common) gets tedious.
Why does the GPU need to handle more than one flavor of Y'CbCr? Why can't the jpeg/PNG/whatever decoder be relied on to convert whatever flavor it uses to the Direct X/OpenGL flavors?
The vast majority of images are jpegs, which are internally 420 YUV, but they get converted to 32 bit RGB for use in apps. Using native YUV formats would save half the memory and rendering bandwidth, speed loading, and provide a tiny quality improvement.
What does he mean by using native YUV formats? Something (I wave my hand) in the rendering pipeline from the JPEG in memory to pixels on the screen?It's been done. There are even GPUs that support this operation natively, so there's no additional overhead.
It's a bit tricky because rendering pipelines composite the final image through many layers of offscreen compositing before the pixel hits the screen.
The core issue is that the offscreen composited layers would still be 32bit textures which is a bigger issue. I would imagine a Skia-based draw list to encode this through the pipeline which could help preserve this perhaps.
Your display uses 3 bytes per pixel. 8 bits for each of the R, G, and B channels. This is known as RGB888. (Ignoring the A or alpha transparency channel for now).
YUV420 uses chroma subsampling, which means the color information is stored at a lower resolution than the brightness information. Groups of 4 pixels will have the same color, but each pixel can have a different brightness. Our eyes are more sensitive to brightness changes than color changes, so this is usually unnoticeable.
This is very advantageous for compression because YUV420 requires 6 bytes per 4 pixels, or 1.5 bytes per pixel, because groups of pixels share a single color value. That's half as many bytes as RGB888.
When you decompress a JPEG, you first get a YUV420 output. Converting from YUV420 to RGB888 doesn't add any information, but it doubles the number of bits used to represent the image because it stores the color value for every individual pixel instead of groups of pixels. This is easier to manipulate in software, but it takes twice as much memory to store and twice as much bandwidth to move around relative to YUV420.
The idea is that if your application can work with YUV420 through the render pipeline and then let a GPU shader do the final conversion to RGB888 within the GPU, you cut your memory and bandwidth requirements in half at the expense of additional code complexity.
Wikipedia is a good source of diagrams and details that explain this further: https://en.wikipedia.org/wiki/YUV#Y%E2%80%B2UV420p_(and_Y%E2...
I can think of several image processing tasks which are more straightforward in a luma/chroma format. Maybe it's because I'm more used to working with the data in that form?
Probably because it's the native input format for every display technology in existence? If you're going to twiddle an image, you might want to do it in the format that it's displayed in.
How would you do even trivial things like color blending in 4:2:0?
That said, this (storing JPEGs in YUV420) is just an optimization and the more images and displays go HDR, the less frequently we'll see YUV JPEGs, though we could see dithered versions, maybe, in 444. That's basically the same thing, but once you discard 420 and 422 you might as well use standard 8-bit RGB and skip the complexity of YCbCr altogether. If you're curious about HDR as I was, though 10-bit is "required", you can dither HDR to 8-bit and not notice the difference unless doing actual colour grading (where you need the extra detail, of course). For obvious reasons, I've never heard of anyone dithering HDR to YUV 420 though, and most computer screens look pretty terrible in when output to a TV as YUV422 or YUV420.
Converting via lookup tables, one for each component, made it very cheap to perform palette-style tricks like tonemapping and smooth color shifts from day to night.
[Edit: it may have actually been CIE Lab, not YUV/YCbCr, because the a&b tended to have narrower ranges. It's been too long!]
I think a twitter account called @JohnCarmackELI5 that just explains all of Johns tweets like I have no idea what he's talking about would be invaluable. The man is obviously bursting with good/interesting stuff to say, but I grok it like 5% of the time.
A great overview and explanation: https://cloudinary.com/blog/time_for_next_gen_codecs_to_deth...
If that isn't an upgrade path, I don't know what is.
There are alternatives. Google optionally uses webp, because it saves them bandwidth and they control the entire chain. And specialized applications like maps sometimes use better suited formats, but if you want to share a picture, that's JPEG.
The same thing happens with audio. The go to format is still MP3 even though it is well outdated. In fact, we had what is close to the perfect lossy audio codec (Opus) since 2012 and support it is pretty much only used when you control the entire chain, like in video games.
So maybe JPEG XL is the perfect image format, but it is still a tough sell when you have something universal that is good enough against it.
The only way I can see it gain traction is if tech giants get together and force it on us. Like they have been doing for web standards. And if there are patents, even that may not work.
Video is different because it uses a huge amount of bandwidth and no formats are "good enough" yet.
However, chroma subsampling is a very primitive form of 2x2 block compression (NB: not related to JPEG's DCT block size). These days GPUs support much better compression natively, with 4x4 blocks, and much fancier modes (ETC1, ASTC). With clever encoding of such textures, it's even possible to achieve compression ratio comparable with JPEG's, while having a straightforward way to convert the compressed file to the compressed texture format.
You can set a time range in Google; "Tools" link under the search bar. DuckDuckGo has this as well. Super useful, especially if there's some recent/prominent news and you want to find things other than this news, or if you're looking specifically for older stuff.
I think you can use "word" (with quotes) to prevent the synonym thing; not entirely sure about that.
I don't know if quotes prevent synonyms, but verbatim mode should.
https://timkadlec.com/remembers/2018-03-22-compressive-image...
Seems to be it?
and it was in IE11 even!
In the case of web browsers it depends how the image is used. An <img> with a large JPEG is probably drawn only once during tile rasterization, and browsers could certainly use the memory savings, so it would probably be a win. But if you had a small JPEG used as a page background and tiled over the whole screen, the memory savings would be small and you'd be wasting power converting the same pixels from YUV to sRGB over and over, so that would likely be a loss.
Further, the transformation from Y'CrCb to gamma-encoded R'G'B' is also a linear operation. Right?
If you want to composite multiple images, do other intermediate processing, or display the image on an arbitrary non-sRGB display, you probably want to convert to a linear space along the way.
I'm not sure I understand where Carmack is coming from here though (am I missing some context? I don't use twitter and these threads are always a huge pain for me to follow especially since Carmack doesn't even bother breaking on full sentences). I don't get how processing in YUV instead of RGB has anything to do with 10bit components for instance.
Also, in my experience most video software deals with YUV natively and only converts as needed. It's probably different in the gaming and image processing world but that's because everything else is RGB and it seems to be a big ask to just tell everybody to convert to YUV.
Besides if quality is of the essence, you will typically store more that 10 bits for internal processing, probably 16 and maybe even floats if you want to have as much range as possible.
I dunno, I won't pretend that I'm smarter than Carmack, but I wish there was a bit more context because it's a bit opaque for me at the moment.
This site (threadreaderapp.com) may be of interest to you. It aggregates threads into a readable column as if it were a single article, here's Carmack's "thread": https://threadreaderapp.com/thread/1400930510671601666.html
Extremely useful for dialogues/conversations on twitter.
For photos I use whatever the camera likes, but for everything else I use 4:4:4.
If you're outputting at a lower resolution or not using a camera then it can be a notable loss of quality.
Most web images are probably scaled down, and in other contexts I bet it's similar. If you're looking at a raw camera shot you're usually only dealing with one at a time, so while simplicity is nice the RAM impact of those will be limited.
I know several imaging apps have the ability to select which subsampling type to use as well.
Seriously though chroma subsampling is not kind on any kind of red shape, it's especially bad against white backgrounds.
It's pretty easy to write a better one since the full resolution image is already available for the Y plane, you can do super resolution.
Most of them. Pretty much the default subsampling for libjpeg(-turbo).
...
> You can do it today, but you need to do the color conversion manually in a shader, which can be a big ask for some devs
Where the perf. really matters, and shaders are involved, doesn't everyone use texture compression formats like ETC1 and PVRTC? https://developer.android.com/guide/playcore/asset-delivery/...
Yes, when they can pay the computational price for compression beforehand. Not going to work, if you had to compress in realtime.
However, I've heard that Netflix was one of the first sponsors of https://github.com/BinomialLLC/basis_universal They are shipping innumerable images to a huge variety of underpowered set top boxes. Totally worth investing in a giant-leap memory optimization.
Many mobile GPUs support an extension that does YUV conversion for you (GL_EXT_YUV_target). Maybe it's of less interest to desktop GPU vendors?
Everyone is currently approaching AGI in what you might call a traditional way. Carmack's way is completely different, which was refreshing.
I think Carmack's technique has the highest chance of reaching AGI. The tweet chain isn't as unrelated as it seems.
You could try asking him. He likes talking about this, especially if you phrase your questions well.
It's the 21st century and brilliant minds are posting thoughts in broken unreadable parts on websites that completely don't work unless you're a religious "the modern web requires javascript REEEEEEE" zealot.
(b) Can you use curl? Maybe hit the API.
(c) The modern web DOES require JavaScript. Sure, I'm not clear on what that has to do with reading three tweets available on a public API, but it's pretty weird to object to a declarative statement of truth.
YUV format is lossy. The one that is twice as bandwidth-efficient as RGB records four times less color difference than full bandwidth YUV or RGB. Add to that that RGB often used with alpha channel and for half the bandwidth you will get twice the artifacts.
The subsampling scheme that is half of RGB is called 4:1:1: https://en.wikipedia.org/wiki/Chroma_subsampling#4:1:1
It is not broadcast quality. 4:2:0 is also not good for today's standards.
4:1:1 is half of RGB but I've never seen anyone ever use it. Pretty much everything in use today (at the consumer level) uses 4:2:0. It's good enough for today's standards (AV1 Main profile only supports 4:0:0 and 4:2:0).
The difference between 4:1:1 and 4:2:0 is that in 4:1:1 the chroma covers a 1x4 line of pixels whereas in 4:2:0 the chroma covers a 2x2 square of pixels.
YUV to RGB conversion (and vice-versa) requires use of conversion coefficients (BT.470, BT.709, etc): https://en.wikipedia.org/wiki/YUV
Different apps/algorithms can choose their own coefficients, so you get a slightly different RGB colors. If you converted RGB back to YUV with a different set of coefficients, you would get a different result.
Assuming infinite precision, sure.
Interestingly however, there's a closely related color space called YCoCg, that actually does achieve lossless coversion to/from RGB (this lossless variant is often referred to as YCoCg-R, the R stands for "reversible"): https://en.wikipedia.org/wiki/YCoCg#The_lifting-based_YCoCg-...
Essentially, rendering JPEGs directly as a GPU texture (since GPUs can already natively use 420 encodings for textures). His point is that this would improve performance, very slightly improve quality, and reduce power usage as there's no colour space conversion.