Color-Diffusion: using diffusion models to colorize black and white images
github.com
github.com
I was really interested in how color was represented in latent space and ran some experiments with VQGAN clip. You can actually do a (not great) colorization of an image by encoding it w/ VQGAN, and using a prompt like "a colorful image of a woman".
Would be fun to experiment with if anyone wants to try, would love to see any results if someone wants to build
A slight nitpick, wouldn't doing diffusion in the latent space be cheaper?
It's easier to get a sense of what's going wrong with a pixel space model though. With latent space, there's always the question of how color is represented in latent space / how entangled it is with other structure / semantics.
Starting in pixel space removed a lot of variables from the equation, but latent diffusion is the obvious next step
Then there is the issue of B&W movies. Using this kind of tech might not give pleasing results as the colors used for sets and outfits were chosen to work well for film contrast and not for story accuracy. That “blue” dress might really be green. (Please, just leave B&W movies the way they are.)
About every completely automated colorized video tends to be pretty bad though. Particularly the YouTube "8k colorized interpolated" kind of low effort channels where they just let them pump out without caring if it's actually any good.
However, just applying a simple filter (or single transform without effort) definitely feels derivative to me.
*for works of fiction these issues vanish, but for any historical or documentary photographs/films, I really hate that I am being lied to.
The last time I checked, “the source is public domain” is not a valid defense against the pro-DRM parts of that law.
That's pretty perverse, got any links for a primer?
But fair point, non-public-domain B&W could be withdrawn from distribution.
(Perhaps it just takes some getting used to. Back when I read a black and white comic for the first time (as a child), I had a hard time figuring out things at first but got used to it at some point.)
For instance, fake blood in B&W was often produced with black liquid. Colorizing it correctly just doesn't make sense. Or a green or blue dress can be chosen because of the way it looks on film, not because it's supposed to BE a green or blue dress.
I'd like to make one exception, though, for They Shall not Grow Old. That was impressive.
One of the nice features of the somewhat old Deoldify colorizer is support for any resolution. It actually does better than photoshops colorization: https://blog.maxg.io/colorizing-infrared-images-with-photosh...
Edit - technically, I suppose, the way Deoldify works is by rendering the color at a low resolution and then applying the filter to a higher resolution using OpenCV. I think the same sub-sampling approach could work here...
But like most diffusion models, they don't generalize very well to resolutions outside of their training dataset
I have to wonder whether it works well with anything else.
but that’s basically the stable diffusion paper (diffusion in latent space plus GAN superres)
https://github.com/TencentARC/T2I-Adapter
i've also seen a controlnet do this.
Wikipedia has a great example image here: https://en.wikipedia.org/wiki/Chroma_subsampling. Most people would say all of them looked fine at 1:1 resolution.
Something I was thinking about after writing the comment is that the model is probably trained on chroma-subsampled images. Digital cameras do it with the bayer filter, and video cameras add 4:2:0 subsampling or similar subsampling as they compress the image. So the AI is probably biased towards "look like this photo was taken with a digital camera" versus "actually reconstruct the colors of the image". What effect this actually has, I don't know!
re. chroma subsampling in training data: this is actually a big problem and a good generative model will absolutely learn to predict chroma subsampled values (or JPEG artifacts even!). you can get around it by applying random downscaling with antialiasing during training.
In a traditional pixel-space (non latent) diffusion model, you noise all the RGB channels and train a Unet to predict the noise at a given timestep.
When colorizing an image, the Unet always "knows" the black and white image (i.e the L channel).
This implementation only adds noise to the color channels, while keeping the L channel constant.
So to train the model, you need a dataset of colored images. They would be converted to LAB, and the color channels would be noised.
You can't train on decolorized images, because the neural network needs to learn how to predict color with a black and white image as context. Without color info, the model can't learn.
I think the parent means with delocorized images used to test the success and guide the training (since they can be readily compared with the colored image they resulted from which would be the perfect result).
Not to use decolorized images alone to train for coloring (which doesn't even make sense).
Black and white film doesn't have one single colour sensitivity. Play around with something like DxO FilmPack sometime (it has excellent measurement-based representations of black and white film stocks).
It's a much more complex problem than it might seem on the surface.
And I think it can't work. But now I am not sure!
The other day I was working on a mono photo to prove a point: that a model (a photographic artist's model!) with very striking pink hair was of little concern to a photographer who worked in black and white only, and might actually present some opportunities for choosing tonal separation that are not present in those with non-tinted hair.
In different circumstances (film and filter) her hair could appear (in black and white) to the viewer as if it was likely brunette or likely blonde, before any local (as opposed to image wide) adjustments were made.
The question you are asking, I think, is could you get the hair colour right based on the impact of those same circumstances on other known objects in the scene.
I think the answer is no, in the main, generally because those objects likely don't survive to make colour comparisons from (and there are known cases where the colourisation of a building has been completely wrong because it had simply been repainted). And also because it's sometimes not even obvious what a structure actually is, without its colour. People who colourise by hand make this mistake too.
But I concede that given that we have to work with contemporary images to have a colour source, randomising the tone curve is the only thing that could work.
An extreme example:
https://www.cabinetmagazine.org/issues/51/archibald.php
https://www.messynessychic.com/2016/05/05/max-factors-clown-...
Colourising old TV footage can only result in a misrepresentation, because the underlying colour is false to have any kind of usable representation on the medium itself.
And this caricatured example underpins the problem with colourisation: contemporary bias is unavoidable, and can be misleading. Can you take a black and white photo of an African-American woman in the 1930s and accurately colour her skin?
You cannot.
AI colorization will, in general, be plausible, not accurate.
So bullshit is the best you're going to get.
Should artists not put their bs in the world? Writers? Musicians? Most of it is made up but plausible to make you feel something subjective.
That's what happens when you are filling in missing info that isn't in your source.
EDIT: Of course, color photography can be “bullshit” rather than accurate in relation to the actual colors of things in the image; as is the case with the red, blue, and green (actual colors of the physical items) uniforms in Star Trek: The Original Series. But, also fairly frequently, lots of not-intentionally-distortive reproductions of skin tones (often most politically sensitive in the US with racially non-White subjects, where there are also plenty of examples of deliberate manipulation.)
But, yes, in general inaccurate color reproduction can be intentionally manipulated with planning to intentionally create appearances in photos that do not exist in reality.
For some it’s more evocative, irregardless of the absolute accuracy.
Having a professional do it for that picture of your great grandad is expensive.
Having a colourisation subreddit do it is probably worse for accuracy.
I think there is a place for this bullshit.
And plausibility is a feauture, not a bug.
There are always many plausibily correct colorizations of an image, which you want the model to be able to capture in order to be versatile.
Many colorization models introduce additional losses (such as discriminator losses) that avoid constraining the model to a single "correct answer" when the solution space is actually considerably larger.
A YCbCr colorspace is directly mapped from RGB, and thus is limited to that gamut.
LAB can encode colors brighter than diffuse white (ala #ffffff), like an outdoor scene in direct sunlight.
Sorta HDR (LAB) vs non-HDR (YCbCr).
This image (https://upload.wikimedia.org/wikipedia/commons/thumb/f/f3/Ex...) is a good demo, left side was processed in LAB, right in YCbCr). Even reduced back down to a jpeg, the left side is obviously more lifelike, since the highlights and tones were preserved until much later in processing pipeline.
> An example of color enhancement using LAB colorspace in Photoshop (CIELAB D50). Left side is enhanced, right side is not. Enhancement is "overdone" to show the effect better.
And per the original upload the “enhancement” demonstrated is linear compression of the a* and b* channels—
https://upload.wikimedia.org/wikipedia/commons/archive/f/f3/...
—the effect a divergence from the likeness of life at least as I’ve experienced it.
Tbh most cost effective would be a conditional GAN though
Then train the model on movies that are color and then turn them black and white.
That way you can train temporal coherence.
24 frames per second * 60 seconds per minute * 90 minute movie length = 129600 frames
If you could get cost to a penny per frame, about $13k? But I'd bet you could easily get it an order of magnitude less in terms of cost. So $1500 or so?
And that's assuming you do 100% of frames and don't have any clever tricks there.
> penny per frame
Where did this come from?
If you wanted to do this at high res, you would definitely use a latent diffusion model. The autoencoder is almost free to run, and reduces the dimensionality of high res images significantly, which makes it a lot cheaper to run the autoregressive diffusion model for multiple steps.
So all this is to say.... I don't think there would be commercial demand to, say, "upgrade" classic movies with color. Those films were shot by cinematographers who were steeped in the black & white medium and made lighting and compositional choices that take greatest advantage of those creative limitations.
There was, and maybe there will be again once we get far enough from the consumer burnout from the absolute deluge of that in, mostly, the 1980s-1990s.
https://en.m.wikipedia.org/wiki/List_of_black-and-white_film...
Here's an example I really enjoyed, of a snowball fight in 1896: https://twitter.com/JoaquimCampa/status/1311391615425093634
Alas there has been serious money in this in the past (VHS and as I understand it US cable TV).
I would not assume that we have more taste now than we did then. (The state of cinema suggests the opposite to me at least.)
when your mom asks you to make a black and white image in color
digital retouchers do this kind of work all day for decades
Eztra happy if it would be possibe to tune denoising, using photos from the same series. Multiframe NLMeans right now is slow and mostly theoretical.
It's the pinnacle of the whole thing: "imagine it for me in a way that conforms to my contemporary expectations".
If you're going to colourise images, have the decency to do it by hand. If possible on a print with brushes.
Edit: didn't think this would be popular. Maybe it's the historical photography nerd in me, but colourising images without effort and thought is like smashing vintage glass windows for the fun of it: cultural vandalism.
Also horrified.
The point I am making is that colourisation is subjective art, and that alone.
Colourisation cannot fail to enforce contemporary biases based on poor understanding of the materials. It will darken or lighten skin inappropriately, and mislead in any number of ways.
Doing it by hand (in photoshop or on a print) acknowledges the inherent bias that is involved in colourisation.
Automating it is banal at best and dangerous at worst; colourised images risk distorting history.
Well, faces still have a certain tint, the sky is mostly blue, the grass green, water is blue, mud pools are brown, the ground too, a lot of historical fabrics are certain inherent colors, known flowers have known colors, brownstones have red/brown color. A lot of it, is just not that subjective.
Besides different color film stock (or camera sensor "color science") can already result in dozens of widely different colorings of the same exactly scene.
Do they? A certain tint?
You cannot accurately colourise skin from photographic film without an _enormous_ amount of knowledge of the taking and processing of the film, and of the lighting and subject.
An AI can't do it any better than a painter. You can't take a scan of a print or a negative and get skin tones right.
Think about how weird the skin tones are from scans of wet-plate photography plates compared to the same process used in antiquity with the aim of producing a carbon print.
Yes. There's just not a single one across all faces - but I wasn't meaning that.
What I mean is, we know the kind of tints a face will have. A face is not suddenly going to be blue or green or poppy red. And by how light a black and white face appears, we can tell quite well if it's a darker one (oilish to brown) or lighter (pinkish towards more pale).
If we get it wrong within a range it's no big deal. Color film stocks would also vary it widely.
Hell, even actual people who met the person we colourise in real life will remember (or even experience in real time) their face's hue somewhat differently each.
This is an enormously important issue.
Black and white films of different technologies and manufacturers and eras actually lighten or darken skin tones. Really very significantly.
And it's not going to be obvious from the final positive, unless there's _extensive_ data with those images about how the photography was done. And there never is.
Editing because I can no longer reply: the question of whether a skin tone is a dark one or a light one has had severe real life impacts on people whose lives are now only represented in photographs. You can't write this off as micromanagement; it's about the ethics of representation.
Is it?
If 2 colour film stocks took the same image of them, it would show their hue a little (or a lot) different.
Even if two different people actually met the same person, they will probably describe their face as slightly different tones from memory. (And let's not even get into different types of color-blindness they could have had).
Hell, a person's hue will even look different to the same person looking at them, in real time, depending on the changes in lighting and the shade at the scene as they talk (e.g. sun behind clouds vs directly sun vs shade vs bulbs).
It's not really "enormously important" to micromanage the (non-existent) exact right brown or right pink.
You shouldn’t write this off as micromanagement; it's about the ethics of representation. It is better to leave the original image uncoloured than to colour it automatically based on some fundamentally ill-informed model.
Hand colouring that image based on individual knowledge (for example that someone could or could not pass as white) is ethically better, if colourised images are needed.
Important nuances of culture and history, important and complex stories of discrimination and survival, are damaged by automatic colourisation by models that have no knowledge of the source of the mono image they are colourising.
Which makes it mostly american baggage. Other places who didn't have that history don't have much of an issue with whether a person is shown this or that exact shade in a photo, as it doesn't change anything, the same way making a white guy a little pinker doesn't change anything.
If anything, an AI trained on a large and diverse dataset is probably going to wind up being much more accurate with regards to skin color than a human colorist would be in most cases.
The problem here isn't whether colorization is done by man or machine; it's just ensuring that colorized photos are identified as such. Which they usually are -- that's not a new problem to be solved.
A diverse data set of black and white images doesn't have any kind of knowledge of the colour sensitivity of the medium in that moment.
What film was it? How was it processed? Is it a scan of a negative or a print? What was the colour of the lighting? Was a particular colour tint filter used on the lens? Was the subject wearing makeup optimised for black and white photography?
The black and white image, standing alone, cannot tell you this, I think. Sure, it might get a bit better at, say, identifying a 1950s TV show. But what is the "correct" accurate colour representation of that scene, when televisual makeup was wildly unnatural in colour?
And the dataset an AI is going to train on should be using original color photos that are then converted to B&W across a wide variety of color curves. So it should be fairly robust to all sorts of film types. So again, I repeat that it's probably going to wind up being more accurate with regard to skin tone than a human (with their aesthetic biases) usually would.
No, indeed. Which is why doing it by hand is more respectful of the notion that it is subjective.
Automatic colourisation is and will be viewed differently, as more "scientific", when it's still absolutely beholden to the same biases and maybe misconceptions that we can't unpick because they come from poor training data.
Finally: "original colour photos" are also a problem. Not only for the part of the history where they don't exist. But also for the part of history (until the early 1960s) when the colour rendition of those photos was false or incomplete. You can get a little closer to understanding what that colour looked like, but it's important to understand that colour emulsions vary in the way they work: it's not black and white film with extra colour sensitivity.
So at best you will be colourising the black and white film to look like the colour film, which is not reality. And there are well-understood problems with correct representation of skin tones with colour film until the mid-eighties.
I can see your point; I just think there's a bigger picture here (pun not intended) that you're not seeing.
Then the solution is to correct that misperception, not deny ourselves a useful tool.
> I can see your point; I just think there's a bigger picture here (pun not intended) that you're not seeing.
My overarching point is that this is a tool like any other. And the idea that "doing it by hand is more respectful of the notion that it is subjective" I will push back on 100%.
There is nothing disrespectful about colorizing a photo, automatically or by hand. But it should always be clearly communicated that it is subjective not objective, whether human or machine.
Again, if someone believes the colorization is somehow "real" or "scientific" because a computer did it, then correct their misbelief. Don't stop using the tool. That's the bigger picture here.
So if you colourise an image of someone who appears to be a light-skinned 1930s African-American with colours that appear to conform to our contemporary understanding of light-skinned Black people of our era, you might be getting it right, of course.
But you might be getting it quite, quite wrong, in a way that matters.
The model doesn't need to touch the lightness channel at all, only predict the noised added to the color channels at train time.
At inference time, we start with a real lightness channel (b/w image), and initialize the color channels to random noise. The model iteratively denoises the color channels while keeping the lightness channel locked.
No, doing it by hand doesn't acknowledge that your interpretation is a fallible interpretation shaped by bias, just like translating a written work (e.g., the Bible, for a noted example where this has been done often without any such acknowledgement being conveyed) by human effort doesn’t do that.
Acknowledging bias in translation of either kind is an entirely separate action, orthogonal to the method of the translation itself.
There's a lot of irony in acknowledging this but not acknowledging that each and everyone of us has their own biases inherent to our perception and experiences.
Like the blue and white dress; we all perceive things differently even on identical images, monitors, screens, etc.
You're imagining this irony to suit your personal requirement that I am wrong or foolish.
But if you look at what I am talking about elsewhere in my comments on this topic, it is human bias that concerns me.
One of the things automated colourisation cannot get right is historical depictions of human skin. In a way that really matters.
Human biases will creep into automatic colourisation because they can't _not_ creep in: the test data cannot fully describe the subject matter so contemporary bias will take over.
One of the areas where this really matters is historical depictions of Black people. Automatically colourising a black and white photo of a Black person's skin will almost certainly get their skin tone wrong in a way that might well very significantly misrepresent their history.
The same is true in mixed cultures all around the world; colourism is as much an issue as racism.
People had different lives based on how their skin tone was perceived. Automatic colourisation will not (cannot!) automatically produce a colour image that fits that experience. Because dark skin can appear light (and light appear dark) depending on complexities of reproduction.
This stuff matters. Hence my position: if you wish to ethically colourise an image, first consider not doing it at all. Second, consider doing it by hand based on real knowledge of the subject (their lived history etc.).
I am fully cognisant of the bias issues here (well, as fully cognisant as a white amateur student of photographic history can be)
What if the "effort" way is less accurate?
But it reflects the fact that an accurate colourisation of a black and white image without access to every possible detail about the scene and processing from the photographer's perspective is impossible.
Black and white film is substantially more complex and varied than people understand. Its sensitivities are complex and vary from processing run to processing run, and people at the time knew of the weaknesses of black and white and often used false colour to get an acceptable rendition.
Colourisation is a form of expression, not a form of recovery.
Accurate colourisation is impossible even in a color photograph. There is no "canonical" film stock that accurately represents all actual real-life colors.
The expectation from colourisation is not an accurate representation of the original colors, but a good application of color based on our knowledge (whether from historical facts a human colorist knows or from training with similar objects and materials a NN did) that matches a realistic representation of the scene.
If a human colourist draws a dress and doesn't know the color of it, nor have they any historical information about what the person depicted wore that day, they're going to take a guess. That's kind of what the NN will do as well.
Yes, recolors can be inaccurate but they can make historical moments feel more alive and connected. At the same time one can imagine the issues of a recolor that is inaccurate and that is troubling with historical photographs.
At the same time I have a bunch of old family photos I'd love to recolorize. Maybe the colors won't be quite right but that's an OK failure mode for family photos!
I'd love to see a version where you can drop just a spot or two of the correct color and let the AI fill it out. My grandmother had stark red hair but most algorithms will color her as a blond. It'd be nice to fix that, using one of the color photos we do have.
But also you have to consider that bias is being introduced in the colour rendition. That causes damage.
For example, you could see a photograph of an African American woman in the 20s or 30s, and your AI would say, this is an African American woman and colour her skin in some way.
But a lighter-skinned-looking African American woman in a pre/early-post-war photo is a challenge. She may have had darker skin -- been unable to "pass" -- and the film simply didn't get that across because of its colour sensitivity.
Or she may actually have been light-skinned and able to "pass" (or wearing makeup that helped).
Automatically colouring that image introduces risks to the reading of history; you can read that woman's entire life completely wrong.
It's also common with photos of men from that era who worked outdoors. Many of them will come across much darker-skinned in photos than they actually would have appeared in real life, because not-readily-visible sun damage can look odd in mono. But if you colourise all those sun-baked people the same way, what happens to those of mixed heritage among them? (A thing that is already rather "airbrushed out" of history.)
Without knowing about the lighting, the material, the processing and the source of the positive (is it a negative scan? was it a good one? or is it a scan of a print?) you cannot make accurate impressions of skin tone.
And given the power and importance of photography in the history of the USA in particular -- photography coincides with and actually helps define the modern unified US self-image -- this is not something to blaze through without care.
This is a far less tricky problem in more homogeneous societies, obviously. But even then, there is this perception from photographs that British women in the 1920s were all deathly pale; colourisation preserves that illusion that actually comes in part from photographic style.
Creating art without actually knowing anything about it is the banal apotheosis application of Diffusion AI. - Artist in me
Using ChatGPT to write essays that are better than anyone could have ever written is the banal apotheosis application of LLMs - Teacher in me
It is already here. Better use, appreciate, and try to understand how it works rather than complaining about it doing a better job. In this instance, for example, the model can be made to generate multiple outputs or even better, generate output based on precise user input.
Colourisation cannot be done accurately from a black and white image without context that is almost always lacking. Hand colouring is less dishonest.
I fully agree that being able to generate an aesthetically pleasing image with an AI that has been optimized to do exactly that is a banal application of creativity.
I do think that AI has incredible potential to make (and become art).
The best AI artists don't just throw art into midjourney, they experiment, create their own secret sauce.
Training models has become an art form in and of itself: ai artists curate incredible datasets and devise recipes for training stunning models. Their workflows span multiple companies / tools / models.
AI just means that the goalposts for creativity are shifting. Boring people will use AI to make boring art, artists will find completely unexpected ways to use the tools we build to create art forms we've never imagined before.
Plenty of people say that about colorization period, which, while I disagree, seems more sensible than your position to me, which just seems to be fetishizing suffering.