MIT Unveils Gen AI Tool That Generates High Res Images 30 Times Faster
hothardware.com
hothardware.com
https://news.mit.edu/2024/ai-generates-high-quality-images-3...
Very glad to see GANs (or GAN-likes) coming back! Also I don't know if the examples were cherrypicked but a lot of the images look more realistic than SD. For example the dog in the snow[0] is severely oversaturated in the SD version giving it that distinct AI look (sometimes referred to as AIslop). Also the landscape scene[1] has a random lens flare in the SD version which makes it look oversaturated as well. The MIT images are much better in this regard.
[0] https://tianweiy.github.io/dmd/images/teaser/teaser2_Page_1_... vs https://tianweiy.github.io/dmd/images/teaser/teaser2_Page_1_...
[1] https://tianweiy.github.io/dmd/images/teaser/teaser2_Page_1_... vs https://tianweiy.github.io/dmd/images/teaser/teaser2_Page_1_...
I don't understand how neutral networks operate, but my layman's guess is that when you sometimes see hands with 5 fingers visible, sometimes 4, sometimes 3, sometimes 2, sometimes 1, and sometimes 0, then it's not immediately apparent that it means that every hand has between 5 and 0 fingers.
Think of it this way, if the AI has ever only seen houses with a maximum of 10 windows in it's entire training set, is it so unthinkable that it sometimes draws a house with 12 windows? that's a "sensible" understanding about how houses have a variable amount of windows. It just doesn't work for fingers.
I'm sure the same issue would arise if humans had other body parts that came in large quantities, but almost everything else is either 1 or 2 like the nose or eyes.
But I think what the grant parent means by conditioning is non-textual conditioning, like ControlNet. This will always be more powerful than trying to describe something by text. Think about the description of a character in a novel vs the movie adaptation.
you do not draw an individual with anomalous hands because you have an ontological model in which "humans normally have five fingers per hand".
Knowing "how the world works" is the appropriate source for subsequent expression of a representation.
Somewhere in the architecture a world model should be formed.
One problem with hands might be that they are comparatively small. The model easily gets the big picture right (head, arms, legs), but hands are so-called high-frequency details and are additionally featured in lots of different positions, which are seldom sufficiently described in the captions of the training data.
Here is the link to the actual tool.
If so I am surprised at how SIMILAR the outputs of both models look in the general layout / framing / composition of the image.
How can for example both fox astronaut images have near identical backdrop of earth and earth alone on the same side of image at same apparent size. Virtually the same shade of deep blue for deep space.
"Lightshow at the Dolomities" output is virtually identical. So similar that they almost look like iPhone versus Samsung Galaxy camera comparison shots of same scene (I am exaggerating of course but almost there).
What results in such a close similarity in outputs? Same training data set?
It is almost like they have the same DNA!
As jzbontar below mentions, the crucial point is that the random noise mask is the same. The diffusion models are trained to turn random noise to an image, and they are deterministic at that - the same noise leads to the same image.
What the authors did here was to find a smart way of training a new model able to "simulate" in a single step what diffusion achieves in many; to do so, they took many triplets of (prompt, noise, image) generated starting from random noise and a (fixed) pretrained stable diffusion checkpoint. The model is trained to replicate the results.
So, it is surprising that this works at all at creating meaningful images, but it would be _really_ surprising (i.e. probably impossible) if it generated meaningful images which were seriously different from the ones it was pretrained with!
Pardon my ignorance ...
Does MIT model then not work as a general text-to-image model to generate novel images based on arbitrary new text prompts that it has not seen before?
My understanding is that this paper by MIT doesn't train any new model from scratch. I takes a pretrained model (e.g. StableDiffusion), which however is trained to do "a small step" only: you fix a number of steps (e.g. 1000 in the MIT paper), and ask the model to predict how to "enhance" an image by a certain step (e.g. of size 1/1000); the constants are adjusted so that, if the model is "perfect", you get from pure white noise to an image in the exact number of steps you set. If I remember correctly how diffusion works, in theory you could set this number to any value, including 1, but in practice you need several hundreds to get a good result, i.e. the original StableDiffusion model is only able to fit a small adjustment.
This new paper shows how to "distil" the original model (in this case, StableDiffusion) into another model. However, unlike typical distillation, which is used to compress a big model into a smaller one, in this case the distilled model is basically the same as the one you start with; but it has been trained with a different objective, namely to transform random noise to the prediction that the original model (StableDiffusion) would make in 1000 steps. To do so, it is trained on a very large amount of triples (text, noise, image). But I don't think you can incorporate into this training procedure other "real" images that are not generated by the model you start with, because you don't have a corresponding noise (abstractly, there is no such concept as "corresponding noise" to a given image, because the relation noise -> image depends on the specific model you start with, and this map is not anywhere near invertible, since not all images can be generated by StableDiffusion, or any other model).
Once the model is trained, you can of course give it a new prompt and, in theory, it should generate something rather similar to what StableDiffusion would generate with the same prompt (hopefully, the example displayed on their web page are not from the training set! Otherwise it would be totally useless). But you should never obtain something "totally different" from what StableDiffusion would give you, so in that sense it's not "general", it is "just" a model that imitates StableDiffusion very well while being much faster. Which is already great of course :-)
Noting the "50" steps they referenced- whereas I've been using LCMs with 3-8 steps for months.
Just trying to understand context of the paper / comparisons etc.
90ms is fast, but like 2-3x speedup (over what I've been seeing) which is still huge, but not 30x
Can someone with more context explain?
Which is in the same broad neighborhood as existing Stable Diffusion 1.5 LCM (and similar to the speedup, compared to SDXL, for the SDXL LCM, Turbo, and Lightning models.)
Seems like these startups with high valuations, are more tenuous even than the startups of years past.
But multiple vendors have already demonstrated viable alternatives.
Section 10 deals with the indemnification they are offering. There are a lot of limitations. It's definitely not terrible, but it's not remotely close to what a business actually wants in an indemnification agreement.
In 10.1 (the indemnification from them to you) does not include "hold harmless", but in 10.2 (the indemnification from you to them) it does. That's not an accident, and those aren't meaningless words ;)
If you get sued and notify OpenAI of the suit, their lawyers take over completely. You must do anything they ask (as long as it is "reasonable"), including sitting for depositions, preserving evidence, participating in the discovery process, etc. If you want to have any involvement beyond being told what to do, you have to hire your own lawyers at your own expense. And at the end of it all, they will come to a settlement with the other side. As long as it is "reasonable", you must sign it.
https://openai.com/policies/service-terms
> If there is a conflict between the Service Terms and your Agreement, the Service Terms will control
> This indemnity does not apply where: (i) Customer or Customer’s End Users knew or should have known the Output was infringing or likely to infringe, (ii) Customer or Customer’s End Users disabled, ignored, or did not use any relevant citation, filtering or safety features or restrictions provided by OpenAI, (iii) Output was modified, transformed, or used in combination with products or services not provided by or on behalf of OpenAI, (iv) Customer or its End Users did not have the right to use the Input or fine-tuning files to generate the allegedly infringing Output, (v) the claim alleges violation of trademark or related rights based on Customer’s or its End Users’ use of Output in trade or commerce, and (vi) the allegedly infringing Output is from content from a Third Party Offering.
If you knew or should have known, you're on your own. If you didn't follow the rules exactly, you're on your own. If you used the output in commerce and they claim a trademark violation, you're on your own.
And lastly, remember from the business terms that OpenAI's lawyers take control and you must do whatever they ask in terms of depositions and discovery? In dealing with all that information, if they see any indication that you no longer qualify for indemnification OpenAI is going to say goodbye, send all the legal bills to you.
It's better than nothing, for sure. But not all that confidence-inspiring.
> While image generators like DALL-E and Meta AI's Imagine can produce extremely impressive results, these groups are highly protective of their technology and jealously guard it from curious public eyes. Meanwhile, you can go read about MIT's findings at this link.[0]