Near-lossless image formats using ultra-fast LZ codecs
richg42.blogspot.com
richg42.blogspot.com
For color images on Windows, D3D 11.0 and newer versions require support of BC7 decompression: https://learn.microsoft.com/en-us/windows/win32/direct3d11/b...
On Linux, some embedded GPUs I targeted supported ASTC decompression: https://en.wikipedia.org/wiki/Adaptive_scalable_texture_comp...
https://github.com/BinomialLLC/basis_universal
Delivering images in a GPU-friendly manner on the web could be a massive memory saver (PNG/JPGs end up being 32 bits-per-pixel in memory, while GPU formats are usually 4bpp or 8bpp) but web standards haven't caught up yet, for now the only way to display these formats on the web is inside a WebGL context. It works, but it's a lot of machinery just to show a picture.
Apparently modern GPUs losslessly compress textures in memory by themselves, but the only info I can find is that this is proprietary AMD/Nvidia/Intel voodoo magic.
Historically lossless compression has also been applied to the depth buffer when rendering, though it isn't done the same way as the lossless framebuffer compression since depth values are floating point with a limited range. You can find more information about how that stuff works online unlike color compression, since people optimize around it - try searching for "hierarchical Z".
Lossless framebuffer compression becomes especially important in cases where you are using MSAA or SSAA, where you might have a framebuffer containing 4 or more "samples" for every rendered pixel. If all of the samples are from the same polygon, it's quite possible that they will have the same value, so it's really profitable to compress them. Then at the end when you do your "resolve" to turn those 4+ MSAA/SSAA samples into a single pixel, that step becomes easier too.
isnt that one called compression just for marketing purposes? Afaik its a caching technique aimed at lowering latency cost of touching Zbuffer. It actually takes more ram than just raw non duplicated data.
Textures can be auto losslessly compressed too, but note it's only useful for textures the CPU/GPU is writing to; you should use ASTC or whatever usual format most of the time.
eg it's briefly discussed here https://developer.apple.com/documentation/metal/textures/opt...
Typically textures are just swizzled.
https://cbloomrants.blogspot.com/2020/06/oodle-texture-bc7pr...
https://cbloomrants.blogspot.com/2020/09/how-oodle-kraken-an...
https://cbloomrants.blogspot.com/2021/02/rate-allocation-in-...
There are other papers and algorithms that are similar in size to Draco, but without the massive overhead, e.g. EXT_meshopt_compression. Even in 90s papers, there are far more competitive algorithms [1]
[0] https://medium.com/box-developer-blog/showdown-3d-model-comp...
[1] For instance, https://graphicsinterface.org/wp-content/uploads/gi1998-4.pd...
I think we'll care more about increase in quality vs. those little cost : disk space is on shared servers / cloud computing, bandwidth is getting better and better, computation power is also increasing all the time and with new methods of doing computation like GPU, Chiplet / SoC / one-task-chip or simply more efficient CPUs etc.
2D images are used for their own sake, not as a compromise because you'd rather show a moving 3D object rendered at 8K but you can't. So obviously this is about the problem of compressing 2D images, it's off-topic to argue that it can't do video or whatever.
We went through that when flash first got popular. There were some good use-cases for sure, but pages that substituted video for images just ended up being far too busy while also being harder to maintain.
For something like that you'd low latency and thus fast decompression.
Orig PNG: 1443667 bytes. Lossy LZ4I = 1433231, delta = -10436 bytes. A more lossy LZ4I = 1344183, delta = -99484 bytes from PNG.
Original PNG: 1310987 bytes
Lossy PNG: 1154386 bytes (90.4% of LZ4i)
Lossy LZ4i: 1276477 bytes
More Lossy PNG: 693278 bytes
More Lossy LZ4i: 818774 bytes (84.7% of LZ4i)
So it's notable that this lossy PNG technique has diminishing returns for LZ4i - it actually becomes more effective as a PNG preprocessor as you get more lossy, relative to LZ4i!Still, there's a decent argument that PNG is no longer Pareto-optimal. There's stuff that gets close to the same compression ratio but decompresses much faster, and there's stuff that compresses much better and is also lossless, e.g. JPEG-XL lossless mode.
However I greatly dislike this trend because I value my storage space. The less storage space a game wastes, the more games I can have installed. I would much rather like to have more games installed than have each game decode bitmaps more efficiently.
I also think it's a false dichotomy to say you can only choose between slow decoding methods with high bandwidth savings and fast decoding methods with low bandwidth savings. zstd has shown that you can have both fast decoding and high compression.
I don't understand why he would be using lz4 when he can have zstd. The speed advantage may look big on paper but nobody cares if your level loads in 3.2 vs 3.1 seconds.
Perhaps my lived experience is too different. I typically keep a small handful of games that I play frequently, like Halo Infinite or Cyberpunk 2077. Both have huge memory footprints, but not much when HD space is measured in terabytes…
0. https://screenrant.com/cod-modern-warfare-file-size-too-big-...
Timeless classics like KOTOR 1&2, Fallout New Vegas, Mass Effect etc take so little space I barely notice the storage.
The entire library of video games ever made before the year 2000 could fit onto a single SD card off the shelf from Best Buy for the cost of a mid-range restaurant dinner for 2.
It's definitely the case, though, that games have essentially topped out on what they need to support a high fidelity experience in terms of "just" pushing more data down the pipe. While new AAA can push a little farther still, there is broad consensus that diminishing returns are here, and...well, the future is going to be in having AI do the details and make the optimizations, and the resulting size from that is most likely going to stay linear with the number of assets, rather than having each asset metastasize further detail.
A principle of least astonishment is being able to memcpy() a structure or mmap() a file, and get the correct numbers in the fields without swapping bytes all the time.
The "problem" is there either way, and I'd rather have it be noticeable immediately, than at an unknown point in the future.
A format that tricks you into thinking you don't need to think about it is a suboptimal one, in my opinion.
Little-endian platforms have won, very few people are targeting MIPS, power or SPARC processors these days. Technically some ARM CPUs are bi-endian, but practically all operating systems people run on them are only supporting little-endian mode. GPUs are little-endian as well.
The overengineering argument would make sense for a specific implementation, but not for the format as a whole, IMHO.
I don’t think many of these legacy BE systems have such a thing. AFAIK most of these systems are headless: supercomputers with Power10, networking equipment, old SPARC servers, etc.
Maybe we are in the transition phase to where little endian IS the least astonishing thing to do :p
You know that sooner or later something else is going to be needed. Eg, HDR.
I tend to have my browser at 50% width split with another window, and the text goes off the screen making it so I have to scroll left/right to read the text.
Without the min-width it is perfectly readable. The images could also benefit from having the width set to 100% so they scale with the page width instead of being clipped on smaller devices/windows.
I have set my browser to default to Reader View Mode for Blogspot, PG's website, DaringFireball and few others.
But the sample images seem to have been originally quite heavily compressed with JPG or something similar, because zooming in gives very visible compression artifacts.
Why not optimize for CO2 emission instead ?
Main problem with the grid in California is when it sets fire to things or windstorms cause an outage. The generation side is holding up pretty well.
Also, renewable energy doesn’t mean infinite/limitless energy. It often means there are more limitations than non-renewable counterparts. Eg: storing energy is still not a solved problem without really toxic (environmentally and politically) materials.
We can estimate the emissions caused by crypto-currency as that is almost entirely CPU-bound by about the same amount (or an easy to estimate average) per unit. The costs of transmission and storage are going to be relatively minor.
For images the compression cost, while the process is similarly CPU-bound, is pretty small unless you are doing something small like compressing to hundreds/thousands of algorithm/parameter combinations and picking the best result (perhaps by some heuristic more complex than final file size). Also transmission and strorage costs are difficult to estimate because for any given image how do you guestimate how many times it will move and be copied?
Having said that, the article's statement that: “[as] bandwidth savings from overly lossy image codecs will become meaningless, [] the CPU/user time and battery or grid energy spent on complex decompression steps will be wasted.” implies that energy use is at least a coincidental consideration here. Though image quality is by far the primary concern, and that despite being subjective is probably an easier metric to measure in this case, with compression time likely to be ahead of energy use too.
Storage at this point is plentiful. Most people can live just fine with 1 TB of it, which is trivially found in both HDD and SSD forms. Using less saves you nothing, as the media is still there.
Bandwidth is also plentiful and available 24/7. It's also very bursty -- you only download your game once, while you may play it a lot. So saving 50% download time makes no difference energy-wise.
What we can control a lot is CPU power usage. We now have 16 core consumer CPUs that are designed to sleep when not in use, so difference in power usage at runtime can be dramatic.
So an algorithm that trades disk space for decoding efficiency is probably the best way to save power.
So on my current connection, that's maybe 2-5 seconds of download time. Meanwhile, the router is on constantly anyway, so whether it's 200 MB or 400MB makes next to zero difference energy-wise.
Okay, say you're on a 10 Mbps connection. That update takes you what, 5 minutes? Yeah, that's an annoyance, but that's time for a coffee or a bathroom break. Most games don't update daily or anywhere close to it, so your download time is still going to be <1% of your play time for most people.
If your hardware/connection is really awful, yeah, maybe downloading the latest DOOM isn't a great experience, but luckily there's no lack of smaller things to play.
Either way, it's completely irrelevant carbon/energy-wise. Your power usage is going to be dominated by a game trying to render 3D graphics at 60 FPS, not by the tiny increase of the power draw of a router during an update download.
https://expressiveaudio.com/blogs/audio-advent/audio-advent-...
https://www.extremetech.com/gaming/189268-digital-game-downl...
https://arstechnica.com/science/2022/05/discs-vs-data-are-we...
for one thing there is the question of how to attribute carbon emissions for all the "middleboxes" that are drawing tiny amounts of power throughout the whole process but could add up to a lot.
I myself live in western Europe and I can't do cloud backups because it would take me about two weeks of 24/7 uploading which then constantly interferes with every other internet connection I want to use and makes things like streaming videos, zoom sessions or online gaming impossible. Game updates usually mean I won't be gaming that evening. And that's already the better connection after switching providers.
The actual optimization here would be fewer updates and games releasing in a finished state.
So, in what plausible scenario would a gamer's CO2 output (which is mostly power use) be significantly influenced by download time?
The vast majority of people spend a lot more time playing than downloading. If your connection is really bad, then you're probably going to buy physical media or play something smaller.
There's no plausible scenario I can think of in which a significant amount of people sits there waiting for an hour for a game to download, on a daily basis, and where the amount of power consumption caused by that is significant when compared to what's consumed by actual game play.
Yeah, I get that giant updates might be annoying to you personally. But the discussion is about CO2.
I have two ADSL connections that can download at 20 Mbits/s with a load balancer.
With my XBOX ONE I don't really have a choice of physical media, even if I have physical media I'd expect to download between 2-40 GB of patches or additional content. It is definitely a hassle but not prohibitive, generally I install a game before I expect to play it and keep up on downloads ahead of time.
For PC games it is similar except most of the games I play I get off Steam I don't even know if physical media is available at all.
At 20 Mbps, 40GB takes a bit over 4 hours to download. Supposing you have an idle power draw of 60W, that's 0.240 KWh. I think that's reasonable. Downloading is almost effortless for a modern computer, and I'll assume that you're not getting any more use of it, so the monitor turned itself off.
Now compare that to a CPU easily having a TDP of 100W, plus a GPU drawing 180W, plus the rest of the computer, plus a monitor, etc, you can easily reach 400W. That'll blow through the power you consumed to download in half an hour.
Also I don't think anyone releases 40GB patches on a daily basis, so if you game for say, 3 hours a day for a week, the costs of the download is already down to 3%.
Which again is what I was getting at: you should optimize the biggest sources of waste first. Anything that makes your actual gameplay more efficient and consume less CPU or CPU resources will decrease your consumption far more noticeably than almost any optimization you can do to downloading stuff.
How if you don’t know where the electricity comes from powering the hardware your image decoder runs on.
The best way to "optimize for CO2 emission" is to just stop living.
But watch out for induced demand (the user turning the graphics settings up because you made the FPS better.)
A super-awesome feature for the web is that it is progressive - loading a beautiful preview with just the first bytes loaded (and the bitstream can be truncated at any point in time - letting the client download only enough bits for the resolution of the viewer - meaning no need for the server to host different sizes of the same image!)
When rendering a web document requires only a small version it needs only request the part it wants. Crappy resolution is desirable on a slow connection, with limited ram or a small screen. If the next customer wants to zoom in on it on his 4K display he can download the entire thing.
Progressieve jpeg can do at most 5 passes and is a lossy format.
It is designed to display fast then improve the resolution. Perhaps there are implementations that can partially load or stop loading it but it wasn't the goal was it?
Imagine a thumbnail. You have some data there. One clicks on it and a larger version is shown using the thumb in stead of starting from scratch, now one zoons in on the image, surely loading different images from scratch every time is not the right approach? in stead we just don't serve the high resolution version that we [obviously] do have.
IIRC Ogg Vorbis also originally supported this for audio but nobody used it.
http://www.libpng.org/pub/png/pngpic2.html
The problem is that it destroys compression rates. PNG filtering and compression depends on being able to make predictions based on adjacent pixels; interlacing ruins that.
Basically, there are lots of ideas to be had (even more for video) writing reliable implementations and finding adoption is hard.