Real-time image editing using latent consistency models
twitter.com
twitter.com
When Diego first showed me this animation, I wasn't completely sure what I was looking at, because I assumed the left and right sides were like composited together or something. But it's a unified screen recording; the right, generated side is keeping pace with the riffing the artist does in the little paint program on the left.
There is no substitute for low latency in creative tools; if you have to sit there holding your breath every time you try something, you aren't just linearly slowed down. There are points that are just too hard to reach in slow, deliberate, 30+ second steps that a classical diffusion generation requires.
When I first heard about consistency, my assumption was that it was just an accelerator. I expected we'd get faster, cheaper versions of the same kinds of interactions with visual models we're used to seeing. The fine hackers at Krea did not take long to prove me wrong!
There is no substitute for real-time when you're doing creative work.
That's why GitHub Copilot works so well; that's why ChatGPT struck a chord with people—it streamed the characters back to you quite fast.
At first, I was skeptical too. I asked myself “what about Photoshop 1.0? They surely couldn't do it in real-time.”. It turns out that even then you needed it. Of course, compute wasn't there to do a simple translation of all rasterized pixel values that form an image within a layer, but there was a trick they did: they showed you the outline that would tell you, the user, where the content _will_ render if you let the mouse go.
You can see the workflow here:
> https://www.youtube.com/watch?v=ftaIzyrMDqE
It applies to general tools too; you can see the same on this MacOS 8 demo (it runs on the browser!):
So did GPT-3, though. ChatGPT (3.5) was a bit faster, but not overly so.
Good point! I agree with it but forgot to mention it: interaction matters.
With GitHub Copilot you are in familiar terrain, your code editor; with ChatGPT, you are talking to it the same way you'd talk to an assistant, via chat/email.
And we, at KREA, don't think it'll be the exception for AI for creativity.
Online consumption shape digital creative work significantly. AI wont help here either.
The moment you try to market this, you need to be prepared for the lawsuit. I’m preparing for one, and all I did was assemble a dataset. This model is built off of work which most people (rightly or wrongly) believe is not yours to sell.
I’m still not sure how I feel about it. I was forced to confront the question a few days ago, and I’ve been in a holding pattern since then. I’m not so much concerned about the lawsuits as getting the big question right. Ethics has a funny way of sneaking up on you in the long run.
At the very least, be prepared for a lengthy, grizzly smear campaign. Two people wrote stories insinuating I somehow profited off of books3. Your crew will be profiting with intent.
One reason I’ve considered bowing out of ML is that I’d rather not be verbally spit on for the rest of eternity. It’s nice to have the support of colleagues, but unless you really care solely about money, you’ll be classified in the same bucket as Zuck: widely respected if successful by the people that matter, but never able to hold a normal relationship again. Most people probably prefer that tradeoff, but go into this with eyes wide open: you will be despised.
The way out is to help train a model on Creative Commons images. I don’t know if there’s enough data. And it’s certainly a bad idea to wait; your only chance of dominating this market is to iterate quickly, which means using existing models. But at this point, lawsuits are table stakes. You need to be prepared for when they happen, not if.
Also, join me in at least one sleepless night pondering the ethics of profiting off of this. Normally people only mention this as a social signal, not because they actually care. But if you sit down and think it through from first principles, the ethics — legality aside - is not at all clear. This also isn’t a case of a Snowmaker startup (https://x.com/snowmaker/status/1696026604030595497?s=61&t=jQ... he notes that this only works when you have the general population on your side. All of those examples are of startups violating the laws that people felt were dumb. Whereas I can tell you from firsthand trauma that copyright enthusiasts are religiously fanatical. Worse, they might be on the right side of the ethics question.
This was the first time in my life that a startup’s ethics gave me pause. Not just yours, but everyone who’s building creative tools off of these models. You’ll face a stiff headwind. Valve, for example, won’t approve any game containing any work generated by your tools. And everyone else is trying to build their own moat.
I’m not saying to consider giving up. I’m saying, really sit down and go through the mental exercise of deciding if this is a battle you want to fight for at least three years legally and five years socially. I’m happy to provide examples of the type of abuse you and your team will face, ranging from sticks and stones’ level insults to people directly calling for criminal liability (jail time). The latter is exceedingly unlikely, but being ostracized by the general public is not.
At the very least, you’ll need to have a solid answer prepared if you start hiring people and candidates ask for your stance. This comment is as much for your team as for you as an investor, since all of you will face these questions together.
I can't reach out to you via twitter as I am not a verified member, so I will reach out via email.
(It’s unfortunate that there’s no way for me to override that setting. The whole reason I made my DMs public is so people can contact me.)
I don't use twitter much. I've emailed you a rather long email already and would prefer that. I emailed shawnpresser@gmail.com
https://www.reddit.com/r/StableDiffusion/comments/17ovb4j/th...
https://www.reddit.com/r/StableDiffusion/comments/17kvpxn/re...
https://www.reddit.com/r/StableDiffusion/comments/17kekea/de...
https://www.reddit.com/r/StableDiffusion/comments/17ecdab/we...
There are more examples on that subreddit.
Given the pace at which features are added to popular automatic1111 repo, this will be added soon allowing you to try this for free on your own machine.
This is in a closed beta for now (while we work on provisioning enough GPU compute) but we're hoping to make this public later.
Is it your own model or is it based on sdxl or something like that?
we are doing tests, will keep you guys posted.
I'll never get tired of seeing the adaptation capabilities of human nature!
(But of course, we'd love to work on this!)
https://www.ikea.com/jp/ja/p/gestalta-artists-dummy-natural-...
Or maybe just pose yourself and upload a picture, trace it?
with LCM-LoRA you can turn models like SDXL into LCMs without need for training and you can add other style LoRAs like the ones you find on civit.ai
in case you're interested, here's the technical report about LCM-LoRA: https://arxiv.org/abs/2311.05556
some links here: - https://arxiv.org/abs/2310.04378 - https://arxiv.org/abs/2311.05556
The process can be viewed as a particle moving in the image space, one step at a time to its final position in the image space, which is the generated image.
consistency model tries to predict the movement trajectory by providing the current position in the image space. Hence, what used to be a step-by-step process becomes a one-step process.
no that wasnt a sufficient explanation for me. what is the prediction method here? why was diffusion necessary in the past? what tradeoffs does this approach have?
While it can do 1-step the output quality looks a ton better with additional steps.
An LCM demands merely 32 A100 GPUs Hours training for 2-step inference, as de- picted in Figure 1. [1]
Now let's look at the caption under Figure 1:
LCMs can be distilled from any pre-trained Stable Diffusion (SD) in only 4,000 training steps (∼32 A100 GPU Hours) for generating high quality 768×768 resolution images in 2∼4 steps or even one step, significantly accelerating text-to-image generation.
The training mentioned in your quote is distillation: it requires a previously trained SD model.
You could have told just by reading the introduction:
We propose a one-stage guided distillation method to efficiently convert a pretrained guided diffusion model into a latent consistency model by solving an augmented PF-ODE.
[1] Section 4.2 in https://arxiv.org/abs/2310.04378
https://arxiv.org/pdf/2303.01469.pdf
the exact neural network used for the prediction method is omitted. apparently many neural networks can be used for this prediction method as long as they fulfill certain requirements.
> why was diffusion necessary in the past?
in the paper, one way to train a consistency model is distilling an existing diffusion model. But it can be trained from scratch too.
"why was it necessary in the past " doesn't bother me that much. Before people know to use molding to make candles, they did it by dipping threads into wax. Why was thread dipping necessary? it's just a stepping stone of technology development.
The stepping stone way of seeing things reminds me a lot to the thesis behind the book “ Why Greatness Cannot Be Planned: The Myth of the Objective” (2015).
From playing around with it a bit locally, LCM is much, much faster, but generally the detail is much lower using the latest SDXL model.
If your prompt is very simple, such as "a boy looking at the moon in a forest" it does pretty well. If your prompt is much more complex and asks for a lot more detail or uses other LoRas, it doesn't do nearly as well as other samplers and generates lower quality, worse matching images. These other samplers take 30-40 steps so it's several times slower.
From what I've seen though, if you use control net, or passing some guideline images in and rely on a simple prompt, like an existing video that you're trying to change the style of, LCM can generate images in near real time like the OP on an RTX4090 and maybe slower cards on smaller/older models
Another benefit is the decreased experimental time. So you can more quickly iterate over seeds to find output you like, and then maybe you can spend some time with other samplers/upscalers with that seed to make the result higher quality
and when that particle has moved to the right location, there is a decoder that converts it into an image. The decoder network knows how to interpret it.
I actually wrote a micro essay on Twitter the other day about the meaning of the classic encoder decoder network. It’s beautiful.
But yeah!
-
For reference: https://x.com/asciidiego/status/1722544108252836119
but first we want to make sure we can get this in the hands of a tight cohort of creatives and see how they use it.
And this is without taking into account WebGPU and other advances in adjacent fields.
This is the Nvidia tool: https://www.nvidia.com/en-us/studio/canvas/
That's why you can only use what they let you paint. In our demo, you can put whatever text you want.
And we want to make it so that you can assign a (custom) label to each shape. So you can type what each shape represents.
not sure if it's possible to just plug-and-play it or if we would need an extra LCM-LoRA for the motion module.
once we have these sort of models producing frames in milliseconds we should be able to do something similar to this demo but with videos.
Based on my general experience with text-to-image stuff, I assume this lack of adherence isn't always the case? Maybe another demo could show what it's like under ideal conditions.
edit: Webcam demo
https://github.com/radames/Real-Time-Latent-Consistency-Mode...
It does however exhibit similar issues but the realtime constraints of live video make it quit interesting.
Perhaps it's because we've been doing generative AI for years now haha
FCB for example.
A compromise: How about near-real time, would that pass under your BS-threshold?
I have to disagree with you.
Of course, people will use it in a weird way—which is exciting.
Anyway it's not about maturity but control over process, you can compare this to AR 3D drawing tools, also very new but already used to make actual art.
Part of me thinks that this is another revolution in graphics the same way Photoshop was where you can work 10x faster. But another one ponders about what happens when we’re dealing with intelligence.
This is a good point, diffusion models are an example of intelligence. Proof of that is that they became ubiquitous in the same year as large language models, thus they are the same.
Artists, intellectuals, journalists, and critics of all sorts have been jailed, beaten, or simply murdered in the past for expressing themselves in ways the government of the time did not like.
What do you think of this technology? What do you envision?
And, like Sam Altman would say, it is going to be net good for society, but that doesn’t mean there aren’t bumps along the way. We will need to learn to navigate them well.