Quite OK Image is now my favorite asset format
nullprogram.com
nullprogram.com
This summer I spent a lot of time researching lossless compression formats as a streaming format for our 3D capture and reconstruction engine. QOI was on our list too.
My biggest criticism about QOI is that in many cases it doesn't compress at all, or even increases size. It works well on the demo images that come with the package but these are carefully selected. It also works well on high-res, high-quality images with lots of single-color background and/or smooth gradients. But give it grainy, noisy, smaller-format images and it completely breaks down.
We ended up sticking with PNG, which btw. has fast open source implementations too [1][2].
[1] https://libspng.org/ [2] https://github.com/richgel999/fpng
The benchmark suite[1] I compiled contains many different types of images. The "industry standard" sets like the Kodak and Tecnick Photos, texture packs and icons are complete. No images have been removed from these sets. Some of the other collections (photos and game screenshots from Wikipedia and those from pngimg.com) are a random sample. I did not remove any images just because they compressed badly.
Also, the fact that some images compress badly is a non issue, if the whole set of assets that you need for your game/app combined have an acceptable compression ratio. If that's not the case, by all means, chose another image format.
On average, for a lot of assets types, QOI usually performs quite ok.
Lossless compression of random noise isn't possible and lossy compression of it requires content aware algorithms (e.g. AI) to get results.
Even if you don't readily see the random noise, it's there. Subtly changing the hues of pixels just slightly enough to be incompressible.
I am imagining something like a "base layer" with averaged brightness + arithmetic coding for differences might do it.
Imagine an image with large blocks of constant colour in 8-bit colour depth has 2 bits of random noise on every pixel. We indeed can't ever compress to less than 2 bits per pixel, but we can still get down to that.
Nice :-)
In QOIP, opcodes themselves are stored in the format, LZ4 or zstd compression
https://github.com/chocolate42/qoipond
QOIR adds: color profile, EXIF, premultiplied alpha, lossy mode, tile support, LZ4 compression, AFAIK it does this while beating the original QOI in terms of speed AND efficiency.
Lots of benchmarks in: https://github.com/nigeltao/qoir
QOIX is my very own version that adds: Greyscale [+alpha] 8-bit image support, lossy 10-bit support (16-bit encoded as 10-bit, intended for elevation maps), so it's really 3 codecs , optional LZ4. https://github.com/AuburnSounds/gamut
The pitch as I can tell is that the specification is simple. So the guy who spends a lot of time writing code based on specifications likes it?
I guess that’s the niche audience. Maybe a better pitch would be “image format specifications should be more like QOI”.
I think this functions as a test case. For a game as simple as a chess UI, PNG would probably be fine unless you're code-golfing on the final output binary size or refusing to use common dependencies. But for some programs (e.g. large video games), preloading all your assets is very common and decoding speed can be crucial. Maybe the assets could even be left in compressed form in memory in order to reduce the system requirements? I'm not sure if this is common or not.
Side note: video game assets are often enormous in size because developers refuse to implement even very basic compression because of the supposed performance impact. Getting the decompression speed up can result in tremendous reduction in disk usage because it allows doing compression without losing performance.
Those games typically don't need lossless image compression. There are much better (and faster) algorithms for lossy texture compression, many of which can even be decoded directly by the GPU. The S3TC family is one popular example.
One interesting commercial product in this space is Oodle Texture: http://www.radgametools.com/oodletexture.htm
More generally, my point about video games was just that you can frequently cut their size by 50-75% just by turning on file compression on the game directory. Developers are obviously missing some easy wins in this area - including huge wins with lossless compression alone.
I tried using it in a WebGL game, but found that a) decoding was too slow, and b) the quality was too low.
PNG is much bigger, so you’d think it wouldn’t be good for a web game, but once an image has been downloaded and cached, the most important factor is the decode-and-upload time. On the web, PNG and JPEG handily beat basisu there.
In a different situation, basisu could have worked out for me. On native rather than web, maybe the decode time would have been fine. With more photographic rather than geometric assets, maybe the quality would have been fine.
1. Larger size on disk
2. Reduced quality at all MIP levels
3. Complexity of each platform wanting its own special format.
Basis fixes 1 (by adding an extra layer of compression) and 3 (by transcoding to almost any format at load time). But it doesn’t fix 2, and it adds another downside -- decompressing and transcoding is relatively slow.
Using a lossy hardware compression (like S3TC) that's natively supported by the hardware is especially a good way to save on GPU RAM! Double the texture resolution for free, essentially.
This claim needs some real world evidence to back it up (and usually it's not about a performance impact, but instead a perceived image quality impact).
IME a lot of care is taken for compressing asset data, and if there would be a chance to reduce the asset size by another few percent without losing too much detail (that's the important part), it would be done. In the end, textures need to end up in memory as one of the GPU compatible hardware compressed texture formats (e.g. BCx) - which is important not only for reducing memory usage, but mainly for increasing texture sampling performance on the GPU (reduced memory bandwidth and better cache locality when reading texture data from GPU memory).
Those hardware-compressed texture formats rule out any of the popular image formats like JPEG, PNG etc... that might have a better overall compression, but are hard to decode on the fly to one of the hardware texture formats (but there are now alternatives like https://github.com/BinomialLLC/basis_universal), but even a generic lossless compressor (even good old zip) on top of BCx is already a good start.
We're talking lossless compression here, so image quality is not the issue.
Fortunately someone else has already done this research. There's a tool for Windows to control the compact.exe behavior for individual folders called CompactGUI: https://github.com/IridiumIO/CompactGUI
They maintain a database of compression results here: https://docs.google.com/spreadsheets/d/14CVXd6PTIYE9XlNpRsxJ...
Reductions in storage use of greater than 50% are so common that they're hardly even worth remarking on. My experience with compressing a bunch of games is that the biggest gains come from compressing bloated asset packs. Hard to know what else could be taking up more than 50% of the storage space in a particular game.
> Side note: video game assets are often enormous in size because developers refuse to implement even very basic compression
Be generous; you even said in your previous sentence you don't know if it's common or not, so how do you know developers refuse to implement basic compression? If it's that easy, I'm sure there are plenty of AAA game studios (including mine) that will hire you to just implement basic compression.
There are enormous tradeoffs to make that are platform dependent; read speeds from external media are _incredibly_ slow but from internal media may be fast. Some assets are stored on disk to mask loading/decompression. Lastly, assets are just enormous these days. A single 4k texture is just shy of 70MB before compression and a modern game is going to be made of a large number of these (likely hundreds of them), and many games are shipping with HDR - these are multiples of the size. That's becore you get to audio, animation, 3d models, or anything else.
My current projects source assets are roughly 300Gb, and our on disk size is 3GB or so. My last project was 5Tb of source assets and 60GB on disk. Of course we're using compression.
He understands why the Xbox doesn’t have the cd slot anymore.
The last of us part 2 and GTA 5 both needed multiple discs on PS4.
Games heavy on prerendered video frequently were multi-disc: Final Fantasy 7 is three discs, 8 is four, 9 is four... Xenogears is multiple discs, as is Chrono Cross. Lots of JRPGs, though not exclusively: Riven was on 5 (!) discs, again a case of a lot of prerendered content.
It was pretty rare on PS2 with DVDs, and as far as I know totally eliminated on PS3 with Blu-ray (Metal Gear Solid 4 has/had an "install swap" where it would copy over from the disc only a chapter at a time, but all coming from one disc).
Then the concept made a comeback by the tail end of the PS4's life, with several games having "play" and "install" discs, though these aren't quite the same experience, as you just use the install/data disc once then put it away. Fittingly one of these PS4 multi-disc games is the remake of Final Fantasy 7 (one that covers only a relatively small portion of the original, to boot).
The thing I said I didn't know whether it was common was keeping assets in compressed form in memory. I admit I don't know much of the specifics of how game rendering works. What I do know something about is the extremely poor compression applied to many video games in their on-disk form. I'm willing to grant that your studio may indeed be an exception to this, but the general principle isn't possible to deny. Just by enabling lossless Windows file compression on a game folder, you can frequently see the size of a game on disk drop by 50% or more, as I discuss in my comment here: https://news.ycombinator.com/item?id=34042164
Surely a lossless format designed specifically for encoding assets would be even more effective and fast than generic Windows file compression!
> QOI has become my default choice for embedded image assets.
Embedded means they are working on stuff with low amounts of storage, computing speed and memory. If you are developing for embedded you are developing for electronics products typicyally. Every byte you can shave off can ultimately increase the return on investment. Using a traditional image format on embedded might not be possible because you don't have the program memory to store a complex decoder. If you decoder is 10 times bigger than the assets it is meant to decode, maybe there is not so much benefit in using it.
The simplest decoder you can go with would be just storing the value of the pixels in a bitmap and teading that out of that array. This however has the downside that you got no compression on the assets at all. If you have very simple color spaces (e.g. 1bit), tiny resolutions and few assets this might be an acceptable choice, but the choice gets worse as you get more assets, higher resolutions more channels or more bits per channel.
That means according to the author there is a space between just storing bitmaps and just using a png decoder where there were no good goto solutions before and they found a good solution with QOI.
What space? I feel like that's an assumption you're making and it might be true or it might not be, but even if you're right that's very vague and I want more specificity.
I wonder if you'd get better compression using a zigzag scan?
PNG is literally just "perform some simple filtering on the image data, then run that through zlib and wrap it with some metadata". It shouldn't come as any great surprise that you can outperform PNG with a newer stream compressor. You'd probably get even better results by using PNG's filters instead of QOI.
Curiously, for some kinds of serialized binary data varints or even zero-packing[2] can be a good enough filter[3].
[1] https://lz4.github.io/lz4/
Did anybody ever try this? It would be quite interesting and does not seem too difficult.
The only thing I could find is https://github.com/catid/Zpng which does not use the normal PNG filtering.
Encoding and decoding speed as well, because qoi is much simpler (simplistic even) it decodes at 3x libpng (which I’d assume is more relevant than the 20x compression for OP, though then encode might be more relevant for other applications but then IO trade offs rear their ugly heads).
It's nice to be able to understand large parts of your stack at a source code level.
I'd like to see more minimal viable X.
Is old timers spent the first part of our career understand most, if not all, of the technology stack. It’s actually really empowering having that level insight. So I can see why people might still crave for those days of simplicity back.
It’s not not a decision I’d personally follow but I do see the attraction in their decision.
If you are artificially or truly constrained in either CPU usage, or maybe just static code size, it seems like it would have some appeal, even if niche.
The benchmarks are misleading. By size, they mostly consist of photographs and other true-color images, which neither compressor handles well. This has the effect of hiding QOI's lackluster performance on synthetic images which PNG does compress well.
In particular, QOI performs dramatically worse than PNG on synthetic images containing any sort of vertical pattern. As a worst-case example, an image consisting of a single line of noise repeated vertically will compress down to basically nothing in PNG, but QOI will fail to compress it at all.
I can see that QOI performs well for what you might use as texture maps for 3D rasterization, for example. Definitely seems to have some applicability to me.
Presumably that’s with stock libpng, which uses zlib. I wonder if anyone tried patching it to use the substantially faster (on x86) libdeflate[1] instead? It doesn’t do streaming, but you shouldn’t really need that for PNG textures.
We're currently fighting to gain 10-20kB in our binary.
But it sounds to me like you're trying to reduce the amount of code on the NVM of an MCU, where space is very limited, but you have slower off-chip (QSPI, SDC, ..) memory where you can store images?
Some images barely compress which means there should be many runs of uncompressed pixels. Even if only a few pixels long, not having a byte of overhead for each should improve size on these pictures. Maybe they just wanted it to win on throughput by having fewer branches.
[1] The sRGB EOTF has a nominal gamma of 2.2 but it is actually a piecewise function with gamma varying from 1.0 to 2.4.
"Linear" means the colors are stored in a state that means you can do linear transformations on them (basic math in other words) without loss. You can't do this with gamma corrected colors - for example as you say, the sRGB EOTF is a piecewise function so by definition it's not a linear transformation.
So yes, when talking about "linear" colors, they still come from some gamut like sRGB or a larger HDR gamut. The linear part means they are safe for use in a renderer's internals.
Also, normal maps should be flagged as linear, since they should not get converted from the sRGB curve upon sampling.
For video game graphics, the ideal color space is "looks good on the median user's television once we apply our gamma curve", which is not actually a real color space. It's more like a statistical distribution of color spaces.
With video game graphics I'm surprised that PC games aren't handled differently. To the extent that TVs aren't garbage (huge assumption) traditional SDR TVs mostly target sRGB. But PC monitors can be almost anything, there are a ton of wide gamut (e.g. P3) screens out there, and saturated games can be retina-searing on them.
But I think kaetemi's comment was talking about the intermediate calculations, which don't usually use sRGB. The "linear" color space used for those doesn't need to have a rigorous definition; even if there's no mathematically correct way to map a video-game-linear color to an sRGB color, the game company just hand-tweaks it until the game looks good.
Can you actually specify a color profile in the driver itself? I haven't had a color managed workflow in forever but I remember only ever being able to tell the OS.
It's the driver that does the linear RGB to display space conversion when the engine swaps the render buffer.
In the case of non-specified probably-sRGB color space I have no idea if it converts from sRGB to your display space, or if it just does nothing because it not specified.
On modern engines, rendering is done in linear. A lot of older engines render in unmanaged maybe-sRGB and look like potato, in which case tagging the textures does nothing.
In general either the final render output is linear or "I don't care" probably sRGB, in the first case it's up to your display driver to do the final conversion to your display color space, in which case colors should look correct, but the game engine may have already limited colors (by some post processing magic) within sRGB boundaries if it's not aware of your display range. Or it may have messed up if it can't make sense if the display range. In the other case, I'm not sure if there's a solid definition, the driver will probably assume the color space is already good.
Normal maps aren't images for all that they have an intuitive visual representation. In particular they don't have colors therefore a color space has no meaning.
ETA: a better analogy is "plain text encoded": it's probably utf8, but it could be windows-1252, windows-1250, or who knows what else.
When loading images into the GPU, you flag them as either sRGB or linear. When sampling from a shader the GPU will convert sRGB to linear (this is a hardware feature because many images are in sRGB space), and not make any changes to textures flagged as linear.
Rendering is done in linear RGB space in properly coded game engines, and the final output is generally also linear (with some post processing magic to stay in boundaries). The final swap will convert it to sRGB or whatever your display profile is set to, this is handled by the display driver as well.
Fortunately I mostly just deal with embedded graphics that do their technically incorrect blending math in unmanaged sRGB/display space. (:
Yes, it’s exactly that, because:
- almost all the time, the display is sRGB, so that’s your gamut;
- almost all the time, you want to do your lighting and compositing in a linear space.
- and sometimes your channels are normals or something, not colors at all, as others have noted.
So it’s linear in the sRGB 0..255 gamut.
It is a bit of a shame that most easy-to-use tools and workflows are limited to sRGB, so it’s really fiddly to support HDR displays and print.
This is becoming less and less of a safe statement. On the one hand "yay HDR"!
On the other hand, we're in for many years of weird bugs. For example recently I've been working on an app that I want to look good in HDR and I tried to share screenshots. Everything looked good to me. People on the other side were complaining about weird colors. Several hours of investigation later I realized that Windows was switching color spaces in my screenshots. Pressing the print screen button gave me a subtly different result than what I was actually seeing and of course I didn't spot it when emailing off the screen shots.
It's just a laptop, nothing special.
Restricting colour to sRGB is going to be the CGA/EGA in a VGA world of the future.
My least fav is when some system i have to deal with creator chose their superior fav obscure format that no one uses and now i have to deal with it too.
It also has the added benefit that many applications still support it natively. I really don't get the appeal of inventing yet another custom file format that has so few objective benefits.
If you go as far as to implement your own decoder anyway, this argument is simply not applicable. Supporting only a subset of the format is perfectly acceptable - see TIFF, for example; not every application supports every possible feature.
The fact is, none of the existing formats is a perfect fit for the embedded use-case here. It makes sense to introduce a new, simple and straightforward format if it addresses certain typical use-cases better than existing formats do.
So might not be best for every situation, but why not have that in your toolkit? Especially if being able to understand all of your stack is important to you.
Soon in an IOT device near you. Ready to be exploited.
This isn't a big justification for it in a texture format, though. A channel swap isn't that expensive to do at load time and the driver can often do it for you during the upload depending on the API you're using. If you really want to load blazing fast your textures should be pre-compressed in hardware formats that you can mmap in, not stored in QOI or PNG.