67 karma · joined January 23, 2023
This is so interesting and seems obvious in retrospect, but super impressive! The code is simple too, going to hack around with this over the weekend :)
I fully agree that being able to generate an aesthetically pleasing image with an AI that has been optimized to do exactly that is a banal application of creativity.
I do think that AI has incredible potential to make (and become art).
The best AI artists don't just throw art into midjourney, they experiment, create their own secret sauce.
Training models has become an art form in and of itself: ai artists curate incredible datasets and devise recipes for training stunning models. Their workflows span multiple companies / tools / models.
AI just means that the goalposts for creativity are shifting. Boring people will use AI to make boring art, artists will find completely unexpected ways to use the tools we build to create art forms we've never imagined before.
The model doesn't need to touch the lightness channel at all, only predict the noised added to the color channels at train time.
At inference time, we start with a real lightness channel (b/w image), and initialize the color channels to random noise. The model iteratively denoises the color channels while keeping the lightness channel locked.
If you wanted to do this at high res, you would definitely use a latent diffusion model. The autoencoder is almost free to run, and reduces the dimensionality of high res images significantly, which makes it a lot cheaper to run the autoregressive diffusion model for multiple steps.
https://github.com/TencentARC/T2I-Adapter
i've also seen a controlnet do this.
It's easier to get a sense of what's going wrong with a pixel space model though. With latent space, there's always the question of how color is represented in latent space / how entangled it is with other structure / semantics.
Starting in pixel space removed a lot of variables from the equation, but latent diffusion is the obvious next step
I was really interested in how color was represented in latent space and ran some experiments with VQGAN clip. You can actually do a (not great) colorization of an image by encoding it w/ VQGAN, and using a prompt like "a colorful image of a woman".
Would be fun to experiment with if anyone wants to try, would love to see any results if someone wants to build
Tbh most cost effective would be a conditional GAN though
But like most diffusion models, they don't generalize very well to resolutions outside of their training dataset
And plausibility is a feauture, not a bug.
There are always many plausibily correct colorizations of an image, which you want the model to be able to capture in order to be versatile.
Many colorization models introduce additional losses (such as discriminator losses) that avoid constraining the model to a single "correct answer" when the solution space is actually considerably larger.
In a traditional pixel-space (non latent) diffusion model, you noise all the RGB channels and train a Unet to predict the noise at a given timestep.
When colorizing an image, the Unet always "knows" the black and white image (i.e the L channel).
This implementation only adds noise to the color channels, while keeping the L channel constant.
So to train the model, you need a dataset of colored images. They would be converted to LAB, and the color channels would be noised.
You can't train on decolorized images, because the neural network needs to learn how to predict color with a black and white image as context. Without color info, the model can't learn.