The human eye is more sensitive to intensity ("average RGB") than to color. You can drop most of the color info from a picture without significantly degrading its viewing quality.
One way to do that is to transform from (r,g,b) to (intensity, chroma1, chroma2) and then downsize the chroma1 and chroma2 channels by half. When you then transform back into (r,g,b), humans can't hardly tell the difference. Whereas if you tried to do that with the intensity channel, the picture would look awful.