I have to wonder whether it works well with anything else.
I have to wonder whether it works well with anything else.
In a traditional pixel-space (non latent) diffusion model, you noise all the RGB channels and train a Unet to predict the noise at a given timestep.
When colorizing an image, the Unet always "knows" the black and white image (i.e the L channel).
This implementation only adds noise to the color channels, while keeping the L channel constant.
So to train the model, you need a dataset of colored images. They would be converted to LAB, and the color channels would be noised.
You can't train on decolorized images, because the neural network needs to learn how to predict color with a black and white image as context. Without color info, the model can't learn.
An extreme example:
https://www.cabinetmagazine.org/issues/51/archibald.php
https://www.messynessychic.com/2016/05/05/max-factors-clown-...
Colourising old TV footage can only result in a misrepresentation, because the underlying colour is false to have any kind of usable representation on the medium itself.
And this caricatured example underpins the problem with colourisation: contemporary bias is unavoidable, and can be misleading. Can you take a black and white photo of an African-American woman in the 1930s and accurately colour her skin?
You cannot.
AI colorization will, in general, be plausible, not accurate.
And plausibility is a feauture, not a bug.
There are always many plausibily correct colorizations of an image, which you want the model to be able to capture in order to be versatile.
Many colorization models introduce additional losses (such as discriminator losses) that avoid constraining the model to a single "correct answer" when the solution space is actually considerably larger.
That's what happens when you are filling in missing info that isn't in your source.
EDIT: Of course, color photography can be “bullshit” rather than accurate in relation to the actual colors of things in the image; as is the case with the red, blue, and green (actual colors of the physical items) uniforms in Star Trek: The Original Series. But, also fairly frequently, lots of not-intentionally-distortive reproductions of skin tones (often most politically sensitive in the US with racially non-White subjects, where there are also plenty of examples of deliberate manipulation.)
But, yes, in general inaccurate color reproduction can be intentionally manipulated with planning to intentionally create appearances in photos that do not exist in reality.
For some it’s more evocative, irregardless of the absolute accuracy.
Having a professional do it for that picture of your great grandad is expensive.
Having a colourisation subreddit do it is probably worse for accuracy.
I think there is a place for this bullshit.
So bullshit is the best you're going to get.
Should artists not put their bs in the world? Writers? Musicians? Most of it is made up but plausible to make you feel something subjective.
I think the parent means with delocorized images used to test the success and guide the training (since they can be readily compared with the colored image they resulted from which would be the perfect result).
Not to use decolorized images alone to train for coloring (which doesn't even make sense).
A YCbCr colorspace is directly mapped from RGB, and thus is limited to that gamut.
LAB can encode colors brighter than diffuse white (ala #ffffff), like an outdoor scene in direct sunlight.
Sorta HDR (LAB) vs non-HDR (YCbCr).
This image (https://upload.wikimedia.org/wikipedia/commons/thumb/f/f3/Ex...) is a good demo, left side was processed in LAB, right in YCbCr). Even reduced back down to a jpeg, the left side is obviously more lifelike, since the highlights and tones were preserved until much later in processing pipeline.
> An example of color enhancement using LAB colorspace in Photoshop (CIELAB D50). Left side is enhanced, right side is not. Enhancement is "overdone" to show the effect better.
And per the original upload the “enhancement” demonstrated is linear compression of the a* and b* channels—
https://upload.wikimedia.org/wikipedia/commons/archive/f/f3/...
—the effect a divergence from the likeness of life at least as I’ve experienced it.
Black and white film doesn't have one single colour sensitivity. Play around with something like DxO FilmPack sometime (it has excellent measurement-based representations of black and white film stocks).
It's a much more complex problem than it might seem on the surface.
And I think it can't work. But now I am not sure!
The other day I was working on a mono photo to prove a point: that a model (a photographic artist's model!) with very striking pink hair was of little concern to a photographer who worked in black and white only, and might actually present some opportunities for choosing tonal separation that are not present in those with non-tinted hair.
In different circumstances (film and filter) her hair could appear (in black and white) to the viewer as if it was likely brunette or likely blonde, before any local (as opposed to image wide) adjustments were made.
The question you are asking, I think, is could you get the hair colour right based on the impact of those same circumstances on other known objects in the scene.
I think the answer is no, in the main, generally because those objects likely don't survive to make colour comparisons from (and there are known cases where the colourisation of a building has been completely wrong because it had simply been repainted). And also because it's sometimes not even obvious what a structure actually is, without its colour. People who colourise by hand make this mistake too.
But I concede that given that we have to work with contemporary images to have a colour source, randomising the tone curve is the only thing that could work.
Wikipedia has a great example image here: https://en.wikipedia.org/wiki/Chroma_subsampling. Most people would say all of them looked fine at 1:1 resolution.
Something I was thinking about after writing the comment is that the model is probably trained on chroma-subsampled images. Digital cameras do it with the bayer filter, and video cameras add 4:2:0 subsampling or similar subsampling as they compress the image. So the AI is probably biased towards "look like this photo was taken with a digital camera" versus "actually reconstruct the colors of the image". What effect this actually has, I don't know!
re. chroma subsampling in training data: this is actually a big problem and a good generative model will absolutely learn to predict chroma subsampled values (or JPEG artifacts even!). you can get around it by applying random downscaling with antialiasing during training.
but that’s basically the stable diffusion paper (diffusion in latent space plus GAN superres)
https://github.com/TencentARC/T2I-Adapter
i've also seen a controlnet do this.