Do grayscale images take less space?
lu.sagebl.eu
lu.sagebl.eu
The title should reflect that. Obviously this is misleading otherwise, there are a million ways of representing grayscale in data.
Sure, you can argue just picking a color channel and making it the grayscale source is not gonna work as expected, but what if we're talking about the data off a pan sensor vs an RGB array? What about a pan sensor created by taking an RGB sensor and removing the filter so it's producing an image with 3x the resolution but in pan?
This is a complicated question.
This should end up being true for any presentation image format (the one you're going to actually show people, not necessarily the one you might do editing work on) optimized for human vision as well.
It's more a property of our eyes than a quirk of JPEG. So I don't find it misleading personally.
Keep in mind, at least in my experience, if you're even remotely thinking about doing anything with grayscale and this question comes up you're probably doing something...interesting (read: multi band remote sensing shenanigans).
Not only that but color and brightness are often correlated, and advanced compression techniques use that property, see "chroma from luma" in AV1.
Without compression, indeed, it can be 3x, or even 4x (for alignment).
This got me thinking about how subtle and slippery the concept of "resolution" is.
The relationship is not strictly 3x; an ideally demosaiced Bayer filtered image still contains some information at the original capture resolution, because R, G, and B are correlated in nontrivial ways. Nyquist is a nice and simple result for the special case where nothing can be predicted about the original function, measurements are assumed to be perfect, and the criterion for reconstruction is "exact" - but in the general case of "recovering information, more is better" even a simple interpolation procedure with tame priors like "luminance exists" might perform so much better than the naive approach that we'd be silly not to use it.
Even the Nyquist frequency itself is less fundamental than it first appears; a less well known result is that a function can also be perfectly recovered if sampled at half the Nyquist frequency, if the slope is captured as well as the value. In other words you can losslessly convert a 2MP image into a 1M value+slope data matrix and back. What resolution is it? Depends on your point of view.
There are some major failure cases involving alternating single pixel width red / black lines. That caused issues with DVDs because the format also used interlaced frames, meaning things like red tail lights at night could end up looking awful.
Luckily AFAIK they've given up on interlacing.
I really think with modern tech they'd be better off trying to compress 4:4:4 instead of immediately throwing out a lot of colour information.
If it's a full scale photograph full of grayscales that's been dithered, then it's a harder comparison to make. Because obviously a ton of useful detail is lost in the dithering, in a way that isn't with lines and letterforms. So the comparison would really have to be with a terrible, blocky JPEG super-compressed file. In theory the JPEG should win.
This page has a few visual examples at the bottom: https://www.cs.unm.edu/~brayer/vision/fourier.html
It converts the image into frequency domain, and reduces precision of the frequency data, with a special case when the data rounds down to zero.
However, whole-image blur crosses the 8x8 block boundaries, so it isn't perfectly aligned with the blur that JPEG uses. Lowering quality setting in JPEG (or using custom quantisation tables) will be more effective.
There's also a fundamental information-theoretic foundation for this — lower frequencies carry less information.
Naively i would assume that even if you have the unnessary the C_b C_r channels in a greyscale image, they are going to compress really well since they have very little information in a greyscale image.
Yeah, half. Provided you don't store a mono sound in a stereo audiofile.
Same thing for grayscale images. If your white is a #FFFFFF instead of a #FF your picture is a color file that coincidentally displays a grayscale image.
You can also typically discard high frequency data in your color channels. So not only are you not storing full precision data, you also don’t need to store roughly a third of the frequency bins. You can think of this like cropping the image but in the frequency domain. The data saving are identical to cropping the image but visually imperceptible.
You can see the sensitivity as part of the ppmtopgm program.
ppmtopgm (https://netpbm.sourceforge.net/doc/ppmtopgm.html)
ppmtopgm reads a PPM as input and produces a PGM as output. The output is a "black and white" rendering of the original image, as in a black and white photograph. The quantization formula ppmtopgm uses is y = .299 r + .587 g + .114 b.
Note the coefficient for blue is 0.114 and green is 0.587.This also has a spoof on Late Lament ( https://youtu.be/VNC54BKv3mc?t=109 ) that gives me a chuckle...
Cold-hearted orb that rules the night
Removes the colors from our sight
Red is gray, and yellow white
But we decide which is right
And which is a quantization error.