OpenAI releases Consistency Model for one-step generation
github.com
github.com
Correct me if I'm wrong, but I didn't see anything about training time or cost. I would be interested to know whether it is more or less expensive to train this model than it is to train Stable Diffusion.
Small models on the scale of a few billions parameters is probably best left to crowd sourcing efforts because almost everyone can do it. Dealing with large models like GPT and engineering entirely new approaches like this one is more expensive and not easily done by the community. So I guess it is the most efficient use of resources on both side.
Like how Whisper unlocks a lot more text data for GPT4.
Abstract of the paper:
> Diffusion models have made significant breakthroughs in image, audio, and video generation, but they depend on an iterative generation process that causes slow sampling speed and caps their potential for real-time applications. To overcome this limitation, we propose consistency models, a new family of generative models that achieve high sample quality without adversarial training. They support fast one-step generation by design, while still allowing for few-step sampling to trade compute for sample quality. They also support zero-shot data editing, like image inpainting, colorization, and super-resolution, without requiring explicit training on these tasks. Consistency models can be trained either as a way to distill pre-trained diffusion models, or as standalone generative models. Through extensive experiments, we demonstrate that they outperform existing distillation techniques for diffusion models in one- and few-step generation. For example, we achieve the new state-of-the-art FID of 3.55 on CIFAR-10 and 6.20 on ImageNet 64x64 for one-step generation. When trained as standalone generative models, consistency models also outperform single-step, non-adversarial generative models on standard benchmarks like CIFAR-10, ImageNet 64x64 and LSUN 256x256.
Came out back in Dec '22
> These models sometimes produce highly unrealistic outputs, particularly when generating images containing human faces. This may stem from ImageNet's emphasis on non-human objects.
I guess we’ll see what other companies do with this research? It would be great to have image generation times that are closer to image search.
[1] https://github.com/openai/consistency_models/blob/main/model...