HNHacker News
TopNewBestAskShowJobs

erwannmillon

67 karma · joined January 23, 2023

submissionscomments
erwannmillon··on Releasing weights for FLUX.1 Krea
hahahah i'm doing well tianpei good to hear from you!
erwannmillon··on Releasing open weights for FLUX.1 Krea
yeah data really is everything that was the number one lesson from this whole project
erwannmillon··on Releasing weights for FLUX.1 Krea
yoo i'm also a researcher on the krea 1 project and happy to answer any questions :)
erwannmillon··on Show HN: Realtime LCM tool
LCMs are actually exactly SD architecture. LCM is initialized from a regular SD unet and finetuned on a new objective. We are already compiling to get to these times. A lot of other people getting sub-100ms times are using fewer inference steps than we do, at a quality tradeoff.
erwannmillon··on Show HN: Realtime LCM tool
i worked on the gpu/infra side of this, so feel free to AMA Ultimately the LCM is just a SD Unet trained with a new objective, so a lot of SD optimizations are transferable to LCMs
erwannmillon··on Show HN: Patterns – Subliminal Messages with ControlNet
I actually worked on this feature at krea, happy to answer any technical questions about how this works / is trained
erwannmillon··on Improve Stable Diffusion Quality with Skip Connection Rescaling (No Training)
In the decoder, the features from the unet blocks get concatenated with features from the encoder layer through 'skip connections'. The paper discusses how rescaling the backbone features (element-wise multiplication by some scalar) before concatenation improves image quality.
erwannmillon··on Improve Stable Diffusion Quality with Skip Connection Rescaling (No Training)
Improve SD image quality and reduce artefacts without any additional training, simply by reweighting skip connections in the decoder stage of a diffusion Unet decoder
erwannmillon··on Fooocus: OSS for image generation by ControlNet author
"Native refiner swap inside one single k-sampler. The advantage is that now the refiner model can reuse the base model's momentum (or ODE's history parameters) collected from k-sampling to achieve more coherent sampling. In Automatic1111's high-res fix and ComfyUI's node system, the base model and refiner use two independent k-samplers, which means the momentum is largely wasted, and the sampling continuity is broken. Fooocus uses its own advanced k-diffusion sampling that ensures seamless, native, and continuous swap in a refiner setup."

This is so interesting and seems obvious in retrospect, but super impressive! The code is simple too, going to hack around with this over the weekend :)

erwannmillon··on Color-Diffusion: using diffusion models to colorize black and white images
I work at krea.ai. We are for sure making art extremely accessible, but we consider it enhancing creativity rather than replacing.

I fully agree that being able to generate an aesthetically pleasing image with an AI that has been optimized to do exactly that is a banal application of creativity.

I do think that AI has incredible potential to make (and become art).

The best AI artists don't just throw art into midjourney, they experiment, create their own secret sauce.

Training models has become an art form in and of itself: ai artists curate incredible datasets and devise recipes for training stunning models. Their workflows span multiple companies / tools / models.

AI just means that the goalposts for creativity are shifting. Boring people will use AI to make boring art, artists will find completely unexpected ways to use the tools we build to create art forms we've never imagined before.

erwannmillon··on Color-Diffusion: using diffusion models to colorize black and white images
So we're using a color space that has two channels dedicated entirely to color, which is the only thing the model needs to learn.

The model doesn't need to touch the lightness channel at all, only predict the noised added to the color channels at train time.

At inference time, we start with a real lightness channel (b/w image), and initialize the color channels to random noise. The model iteratively denoises the color channels while keeping the lightness channel locked.

erwannmillon··on Color-Diffusion: using diffusion models to colorize black and white images
Think inference time was on the order of 4-5seconds per image on a v100, which you can rent for like .80 cents an hour, though you can get way better gpus like a100s for ~1.1 usd/h now. But ofc this is at 64px res in pixel space.

If you wanted to do this at high res, you would definitely use a latent diffusion model. The autoencoder is almost free to run, and reduces the dimensionality of high res images significantly, which makes it a lot cheaper to run the autoregressive diffusion model for multiple steps.

erwannmillon··on Color-Diffusion: using diffusion models to colorize black and white images
Yeah, if you have a high res image, you can get color info at super low-res and then regenerate the colors at high res with another model. (though this isn't an efficient approach at all)

https://github.com/TencentARC/T2I-Adapter

i've also seen a controlnet do this.

erwannmillon··on Color-Diffusion: using diffusion models to colorize black and white images
Depends, given the low res, the 3x64x64 pixel space image is smaller than the latents you would get from encoding a higher-res image with models like VQGAN or the stablediff VAE at their native resolutions.

It's easier to get a sense of what's going wrong with a pixel space model though. With latent space, there's always the question of how color is represented in latent space / how entangled it is with other structure / semantics.

Starting in pixel space removed a lot of variables from the equation, but latent diffusion is the obvious next step

erwannmillon··on Color-Diffusion: using diffusion models to colorize black and white images
Think the final training run was only a couple hours on a Colab V100
erwannmillon··on Color-Diffusion: using diffusion models to colorize black and white images
Took a lot of failed experiments, the model would keep converging to greyscale / sepia images. Think one of the ways I fixed was by adding an greyscale encoder to the arch. Used its output embedding as additional conditioning. Can't remember if I only added it to the Unet input or injected it during various stages of the unet down pass.
erwannmillon··on Color-Diffusion: using diffusion models to colorize black and white images
Btw, I did this in pixel space for simplicity, cool animations, and compute costs. Would be really interesting to do this as an LDM (though of course you can't really do the LAB color space thing, unless you maybe train an AE specifically for that color space. )

I was really interested in how color was represented in latent space and ran some experiments with VQGAN clip. You can actually do a (not great) colorization of an image by encoding it w/ VQGAN, and using a prompt like "a colorful image of a woman".

Would be fun to experiment with if anyone wants to try, would love to see any results if someone wants to build

erwannmillon··on Color-Diffusion: using diffusion models to colorize black and white images
Fair enough. Honestly this was just a fun side project. I actually coded this up last october when I was doing a deep dive to learn about diffusion models, and saw that no one had ever applied them to colorization. This was just a fun opportunity to build a project that no one had done before
erwannmillon··on Color-Diffusion: using diffusion models to colorize black and white images
temporal coherence is def an issue with these types of models, though I haven't tested it out with ColorDiffusion. Assuming you're not doing anything autoregressive (from frame to frame) to do temporal coherence, you can also parallelize the colorization of each frame, which would affect cost.

Tbh most cost effective would be a conditional GAN though

erwannmillon··on Color-Diffusion: using diffusion models to colorize black and white images
Technically yes, the encoder and unet are convolutional and support arbitrary input sizes, but the model was trained at 64x64px bc of compute limitations. You could probably resume the training from a 64x64 resolution checkpoint and train at a higher resolution.

But like most diffusion models, they don't generalize very well to resolutions outside of their training dataset

erwannmillon··on Color-Diffusion: using diffusion models to colorize black and white images
Yeah, the model is racist for sure. That's a limitation of the dataset though (celeb A is not known for its diversity, but it was easy for me to work with, I trained this model on Colab)

And plausibility is a feauture, not a bug.

There are always many plausibily correct colorizations of an image, which you want the model to be able to capture in order to be versatile.

Many colorization models introduce additional losses (such as discriminator losses) that avoid constraining the model to a single "correct answer" when the solution space is actually considerably larger.

erwannmillon··on Color-Diffusion: using diffusion models to colorize black and white images
hahaha it reminded me of some "zoom and enhance" stuff when I was making the animations
erwannmillon··on Color-Diffusion: using diffusion models to colorize black and white images
You can do this with spatial palette t2i or controlnet. Give a super lores spatial palette as conditioning like this: https://camo.githubusercontent.com/8e488996fd309165fb065b0cd...

https://github.com/TencentARC/T2I-Adapter

erwannmillon··on Color-Diffusion: using diffusion models to colorize black and white images
touché, nevertheless, colors go brrrrrrrr
erwannmillon··on Color-Diffusion: using diffusion models to colorize black and white images
basically the training works as follows: Take a color image in RGB. Convert it to LAB. This is an alternative color space where the first channel is a greyscale image, and two channels that represent the color information.

In a traditional pixel-space (non latent) diffusion model, you noise all the RGB channels and train a Unet to predict the noise at a given timestep.

When colorizing an image, the Unet always "knows" the black and white image (i.e the L channel).

This implementation only adds noise to the color channels, while keeping the L channel constant.

So to train the model, you need a dataset of colored images. They would be converted to LAB, and the color channels would be noised.

You can't train on decolorized images, because the neural network needs to learn how to predict color with a black and white image as context. Without color info, the model can't learn.

erwannmillon··on Color-Diffusion: using diffusion models to colorize black and white images
trained on celebA, so no, but you could for sure train this on a more varied dataset