DeepFloyd IF: open-source text-to-image model
github.com
github.com
(via https://news.ycombinator.com/item?id=35743727, but we've merged that thread into this earlier one)
Colab Notebook for running the model based on the diffusers library: https://colab.research.google.com/github/huggingface/noteboo...
Hugging Face Space for testing the model: https://huggingface.co/spaces/DeepFloyd/IF
Note that the model is substantially more compute-intensive than Stable Diffusion, so it may be slower even though that space is running on an A100.
At that point it's probably better to write a guide on how to set up a VM with a A100 easily instead of trying to fit it into a Colab GPU.
Any more specifics on this? Sampling has gotten better too. My primary concern is the amount of memory necessary for a generation (batch size = 1, or I guess "2" using classifier free guidance).
Similar to OpenAI's cascaded diffusion models and GLIDE, you can presumably run the models in sequence, unloading earlier models from memory to make room for the 2nd and 3rd stage models.
Right now, I really only need 256px resolution. So, will I be able to fit the first stage (64px) model in memory on its own with my 12 GB RTX 3060? What about the 2nd stage (64px -> 256px)? 3rd stage?
It is able to put people on the left/right and put the correct t-shirts and facial expressions on each one. This is compared to mj which just mixes together a soup of every word you use and plops it out into the image. Huge MJ fan of course, it's amazing, but having compositional power is another step up.
i.e. if you say "a red cube on a green sphere" in DeepFloyd, you will get it. If you say that in MJ, you won't. That means you have more power to compose the image you want with this tool.
MJ and stable also make clear that their models don't understand language like humans do.
I believe the claim here is "and we would like them to".
https://twitter.com/eb_french/status/1651365078137200640
BTW I am a HUGE fan of MJ, and attend the office hours, and have done 35k+ images there. So you may have misinterpreted how much of a supporter of it I am.
a single (green sphere), with a single (red cube), balancing ((on top))
0/16 images have a red cube on a green sphere.
Just look at their cherry picks in this discord... https://discord.com/invite/pxewcvSvNx . It's overfitted on images with copyright (afghan girl) and doesn't show more "compositional power" at all most of the time ignoring half of the prompt.
Being "not one shot" for most nontrivial prompts is a failure of current t2i models, its what they all strive for and its what DF supposedly does a lot better. And, while its possible to spin things pretty hard when people can't bang on it themselves, I think the indication is that it is, in fact, a major leap forward from the best current consumer-available t2i models (it looks pretty comparable to Google Imagen – a little bit worse benchmark scores – which is unsurprising since it seems to be an implementation of exactly the architecture described in Google's Imagen paper.
> It’s overfitted on images with copyright (afghan girl)
It’s…not, though. Sure, the picture with a prompt which is suggestive of that (down to even specifying the same film type) gives off a vibe that completely feels, if you haven’t recently looked at the famous picture but are familiar with it, like a “cleaned up” version of that picture, so you might intuitively feel its from overfitting, that it is basically reproducing the original image with slight variations.
Look at the two pictures side-by-side and there is basically nothing similar about them except exactly the things specified in the prompt, and pretty much every aspect of the way that those elements of the prompt is interpreted in the DF image is unlike the other image.
It's not a major leap not even a small one because it's exactly like imagen. It's stability giving some Ukrainian refugees compute time to train "their" model for publicity. It's about the whom and not what as it should be.
I "feel" nothing I am telling it how it is. Look at the afghan girl example again it. Close up portrait, same clothing, same comp, expressive eyes... and most important burn in like every other overfitted image in diffusion networks.
You guys all want it to be something special and I get it, new content, new shiny toy but it's neither a good architecture nor a good implementation.
No, I’m not affiliated with StabilityAI
> It’s not a major leap not even a small one because it’s exactly like imagen.
I would agree, if imagen was a “consumer-available t2i model”. What’s available is a research paper with demo images from Google. The model itself is locked up inside Google, notionally because they haven’t solved filtering issues with it.
> Look at the afghan girl example again it. Close up portrait, same clothing, same comp, expressive eyes…
You look at it again, literally none of those things are the same: its not the same clothing (the material and color of the head scarf is different, the headscarf is the only visible clothing in the DF image, whereas that is not the case in the famous image), the condition of the head scarf is different, the hair color is different, the hair style is different, the hair texture is different, the face shape is different, the individual facial features are different, the eye color is much more brown in the DF image, the facial expression is different, the DF image has lipstick and eyeshadow, the famous image has a dirty face and no makeup, the headscarf is worn differently in the two images, the background is different, the lighting is different, and the faces are framed differently.
The similarities are (1) its a close up portrait, (2) a general ethnic similarity, and (3) they are both wearing a red (though very different red) head scarf, (4) and they are both looking straight into the camera. (2)-(4) are explicitly prompted, (1) is strongly implied in the prompt addressing nothing that isn’t related to the face/head. This isn’t “overfitting on a copyright image” its getting what you prompt, with no other similarity to the existing image.
> You guys all want it to be something special and I get it,
I’m actually kind of annoyed, because I’ve been collecting tooling, checkpoints, and other support for, and spending quite a bit of time getting proficient in dealing with the quirks of, Stable Diffusion. But, that’s life.
> it’s neither a good architecture nor a good implementation.
I’d be interested in hearing your specific criticism of the architecture and implementation, but hopefully its more grounded in fact than your criticism of the one image...
But in the end, people want pretty pictures. So is a complicated situation.
Midjourney does much better overall. Composition is neat, but MJ is so incredibly far ahead in terms of quality of output, it honestly doesn't matter if you have to go and do composition manually (and with new AI based tools, that's easier than ever too. Do a bad cut and paste job then infill your way back to a coherent image)
But if I actually wanted a useful picture, I could work with what MJ gave me despite having minimal image editing skills. The DeepFloyd result looks like it's a 8-12 months behind what MJ gave and wouldn't be salvageable.
https://twitter.com/eb_french/status/1651584746089218049
it wasn't.
Again, I don't know why everyone is so defensive. I love MJ. There's nothing wrong with admitting that other models might do certain things better. We all can use any model we want.
It's not ad hominem at all... mj isn't as good at certain types of composition as others. I don't get why people are pretending that isn't the case. I want everyone to have great models and IF is part of that progress. Perhaps calling it "word soup" was offensive? This isn't your religion, though, it's just a model. Listening in on the MJ office hours they're the farthest thing you can be from dogmatic or arrogant. They want to improve as we all do. I personally am just really inspired that everyone can advance together!
Also see downthread - the first 32 images I generated attempting to reproduce the claim that "actually MJ can do this" all failed. The person who challenged me then ignored it. This isn't really up for debate until someone sends a seed where mj can do the cube + sphere thing well.
Many times you'll have no other choice but to use a diffusion model with img2img.
I agree with OP though, the market has spoken and the vast majority of people use prompts hardly more nuanced than a 90s Mad Magazine book of Mad Libs.
Overall this feels like trying to get ChatGPT to do math: just let ChatGPT offload math to Wolfram.
Similarly I'd rather just offload the composition. Now we even have SAM which will happily pick out the parts of the image you want to compose
It's interesting to see what IF can do in terms of composition, text rendering etc, it's very promising if aesthetically pleasing images can be achieved via fine-tuning (the same happened with SD... current publicly fine-tuned models can achieve much higher levels of quality and cohesion than the base models, here's the prompt in an SD2.1 based model: https://imgur.com/a/ELGMSmV ).
Of course fine-tuning IF is likely more challenging, as both the two first stages and the 4x SD upscaler might need to be fine-tuned...
I really hope that isn't actually possible.
What's really required is semantic composition. Making subjects meaningfully and predictably interact, or combining them together. And also the coherence of the overall stitched picture, so you don't end up with several different perspective planes.
1. A stained glass picture of a woman in a library with a raven on her shoulder with a key in its mouth
2. An oil painting of a man in a factory looking at a cat wearing a top hat
3. A digital art picture of a child riding a llama with a bell on its tail through a desert
4. A 3D render of an astronaut in space holding a fox wearing lipstick
5. Pixel art of a farmer in a cathedral holding a red basketball
And followed up with this article when he won the bet: https://astralcodexten.substack.com/p/i-won-my-three-year-ai...
"2. All persons obtaining a copy or substantial portion of the Software, a modified version of the Software (or substantial portion thereof), or a derivative work based upon this Software (or substantial portion thereof) must not delete, remove, disable, diminish, or circumvent any inference filters or inference filter mechanisms in the Software, or any portion of the Software that implements any such filters or filter mechanisms."
It can be modified. That just says it can't be modified to bypass their filters.
To remove filters.
"Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions:"
[0] https://github.com/deep-floyd/IF/blob/main/LICENSE-MODEL#L54
I'm starting to get irritated by all these 'non-commercial' licensed models, though, because there is no such thing as a non-commercial license. In copyright law, merely having the work in question is considered a commercial benefit. So you need to specify every single act you think is 'non-commercial', and users of the license have to read and understand that. Even Creative Commons' NC clause only specifies one; they say that filesharing is not commercial. So it's just a fancy covenant not to sue BitTorrent users.
And then there's LLaMA, whose model weights were only ever shared privately with other researchers. Everyone using LLaMA publicly is likely pirating it. Actual weights-available or Free models already exist, such as BLOOM, Dolly, StableLM[0], Pythia, GPT-J, GPT-NeoX, and CerebrasGPT.
[0] Untuned only; the instruction-tuned models are frustratingly CC-BY-NC-SA because apparently nobody made an open dataset for instruction tuning.
[1] Insamuch as an AI model trained on copyrighted data can even be considered Free.
Theoretically large companies or rich people might be able to make a licensing agreement.
I thought that at first, but I think it only prohibits commercial use that breaks regional copyright or privacy laws.
EDIT: That said, it’s unambiguously not open source.
Right, but to what end? The only reason the licensor should care one way or another is the licensor being held liable for what folks do with the software, in which case...
> it also means, and this is probably more important, that in some juridictions, the model licensor has a better defense against liability for contributory infringement if the licensee infringes.
Do hardware stores need to demand "thou shalt not use this tool to kill people" to their customers to avoid liability for axe murders under such jurisdictions? Or car manufacturers needing to specify "you will not use this product to run over schoolchildren at crosswalks"?
Like, I'm sure such jurisdictions exist, but I somehow doubt license terms in an EULA nobody (except for us nerds) will ever read would be sufficient in such a kangaroo court.
(EDIT: also, I'm pretty sure the standard warranty disclaimer in your average FOSS license already covers this, without making the software nonfree in the process)
Generally, not, because vicarious liability for battery and wrongful death doesn’t work like, e.g., contributory copyright infringement.
> I'm pretty sure the standard warranty disclaimer in your average FOSS license already covers this
No, warranty disclaimers don’t cover this, because (1) its not a warranty issue, and (2) disclaimers, if they have legal effect at all, effect liability the disclaiming party would otherwise have to the party accepting the disclaimer, not liability the disclaiming party would have to third parties.
Judging by youtube-dl, it seems like it does work that way, at least in my jurisdiction; I guess we'll see if the RIAA doubles down on trying to wipe it from the face of the Earth, but considering there hasn't been much noise, I wouldn't count on it. Also, to my adjacent point, I highly doubt the RIAA would've refrained from attempting to take down youtube-dl even if youtube-dl's license prohibited its users from circumventing DRM with it.
There are current open cases of people claiming “harm” for misinformation spouted by ChatGPT where it just makes up facts to satisfy user prompts.
There are current open cases where people are claiming “copyright violation” due to diffusion models satisfying user prompts.
AFAICT none of these cases are against the users who are prompting the models.
There are also problematic restrictions on your ability to modify the software under clause 2(c). And nor do you have the right to sublicence, it's not clear to me what rights somebody has if you give them a copy.
Your second sentence contradicts the first. Prohibiting commercial use and prohibiting modification are each in and of themselves mutually exclusive being being "technically open source" (let alone both at the same time).
Are people suggesting that "look at the code but don't touch" actually fits what some people think of as open source?
In theory, it is also smarter at learning from its training data.
Not really, it's a cascaded diffusion model conditioned on the T5 encoder, there is nothing really in common, unless you mean that using a diffusion model is "SD style".
Yeah, it looks exactly (architecturally) like Imagen.
Google would be running circles around everyone in Generative AI (maybe OpenAI would still have a better core LLM, maybe, but portfolio-wise) if they simply had the ability to cross the gap between building technologies and writing up research papers on them and actually releasing products.
OSError: DeepFloyd/IF-I-IF-v1.0 is not a local folder and is not a valid model identifier listed on 'https://huggingface.co/models' If this is a private repository, make sure to pass a token having permission to this repo with `use_auth_token` or log in with `huggingface-cli login` and pass `use_auth_token=True`.
From that - I would strongly suspect the answer to your question to be yes.
EDIT: According to the figure in the Imagen paper FL33TW00D's response referred me to, it looks like the text encoder size is the biggest factor in the improved model performance all-around.
> a photograph of raccoon in the woods holding a sign that says "I will eat your trash"
A photograph of an English professor in the woods holding a sign that says "I will eat your trash"
My first thought on seeing "Floyd" and "IF" together. It looks like a Pink Floyd reference from the About page on https://deepfloyd.ai/ though.
Yeah good luck figuring out that a particular logo was generated with this particular model. And if someone does good luck doing anything about it.
With this amount of fear one wouldn't dare to cross a road without three layers of bubble wrap, plus written authorisation from a lawyer plus a feasibility study from a traffic engineer.
Explain it to me what will happen with the generated logo at an acquisition or due diligence.
PS: not an AI apologist, just pointing the irony. Feels like those fan sonic characters "original content do not steal".
Admittedly it was one of several hundred I had it spit out for me. But the design was completely original and caught me by surprise.
I tried making modifications but everyone kept telling me they prefer the original exactly as MidJourney made it!
> Hands
good god it solves the two biggest meme issues with image models in one go. Will this be the new state of the art every other model is compared to?
So, this is going to have new different issues. Since it's similar to Imagen, it probably can't handle long complex prompts as well, since they developed Parti afterward.
Here's my question: are there any image models where, if you prompt "1+1", you get an image showing "3"?
Well, yeah, its a bigger set of models (particular the language model) that takes more resources (both to train and for inference.) That’s the tradeoff.
> Here’s my question: are there any image models where, if you prompt “1+1”, you get an image showing “3”?
You want a t2i model that does arithmetic in the prompt, translates to it to “text displaying the number <result>”, but, also does the arithmetic wrong?
Yeah, I don’t think that combination of features is in any existing model or, really, in any of the datasets used for evaluation, or otherwise on anyone’s roadmap.
"Actually thinking about your prompt" is a necessary part of being able to make the prompts natural language instead of a long list of fantasy google image search terms.
Useful example being "my bedroom but in a new color", but some things I've typed into Midjourney that don't work include "a really long guinea pig" (you get a regular size one), "world's best coffee" (the coffee cup gets a world on it), etc. It's just too literal.
And yes, preprocessing with an LLM could do this.
There are multiple ways to speed up the inference time and lower the memory consumption even more with diffusers. To do so, please have a look at the Diffusers docs:
Optimizing for inference time [1]
Optimizing for low memory during inference [2]
[1] https://huggingface.co/docs/diffusers/api/pipelines/if#optim...[2] https://huggingface.co/docs/diffusers/api/pipelines/if#optim...
Can anyone explain why it needs so much ram in the first place though? 4.3B is only ~9GB at 16bit (I'm not as familiar with image models).
I'm really happy to see that fits under 24GB - that's what I consider the limit for being able to run on "consumer hardware".
The T5-XXL text encoder is really large, also we do not quantize the UNets, the UNet outputs 8-bit pixels, so quantizing the UNet to that precision will create pretty bad outputs.
1. (11B) T5-XXL text encoder [1]
2. (4.3B) Stage 1 UNet
3. (1.3B) Stage 2 upscaler (64x64 -> 256x256)
4. (?B) Stage 3 upscaler (256x256 -> 1024x1024)
Resolution numbers could be off though. Also the third stage can apparently use the existing stable diffusion x4, or a new upscaler that they aren't releasing yet (ever?).
> Once these are quantized (I assume they can be)
Based on the success of LLaMA 4bit quantization, I believe the text encoder could be. As for the other modules, I'm not sure.
edit: the text encoder is 11B, not 4.5B as I initially wrote.
The architecture here looks different, but the code is licensed in a way which still makes downstream optimization and redistribution possible, so maybe there will be something there.
> Aristotle in ancient greek clothes. Toga. New york, rain, film noir, fog, art deco, neon lights, blade runner sci fi
Seems to be holding up recently well with the first promt. Second was only OK.
>ChatGPT explain this like I'm 5
I'm also very happy for the release of the two upscaler, I can use them to upscale to result of my small 64x64 DDIM models (maybe with some finetuning).
I believe the current SOTA test for NLVR is VQAv2[0] or GQA[1].
0: https://visualqa.org/ 1: https://arxiv.org/pdf/1902.09506.pdf
For a more-robust-but-hard-to-run model, you can use BLIP2: https://huggingface.co/Salesforce/blip2-opt-2.7b
Image: https://i.imgur.com/husplYZ.png
Output: "a white horse with a sign that says rexel's in space, pixelperfect, inspired by Paul Kelpe, official simpsons movie artwork, alternate album cover, in style of nanospace, by Apelles, pickles, pespective, pop surrealism, ingame, in a space cadet outfit, sifi"
Hope not. This is a worse license.
To make it concrete, one could argue that bank robbers rob banks even though it is illegal, so why have a law against it since law abiding people aren't going to rob banks. Does anyone really think we should remove such laws?
At a later point the model will be renamed "StableIf", and released with a similar license to StableDiffusion.
DeepFloyd IF is a state-of-the-art text-to-image model released on a non-commercial, research-permissible license that provides an opportunity for research labs to examine and experiment with advanced text-to-image generation approaches. In line with other Stability AI models, Stability AI intends to release a DeepFloyd IF model fully open source at a future date.
DeepFloyd IF is based on Google's Imagen model, which has two key differences from Stable Diffusion: (1) it denoises in pixel space instead of a compressed latent space, and (2) it uses a 10x larger pretrained text encoder (T5-XXL-1.1) compared to SD's CLIP encoder. (1) allows it to better render high-frequency details and text, and (2) allows it to understand complex prompts much better. These improvements come at the cost of multiple times more memory usage and compute requirements compared to SD, though.
In terms of "will it replace SD?"—in the short term I think yes. But I still think latent diffusion models are the future. For example, Stability is gearing up to release Stable Diffusion XL right now, a larger version of the original SD that does higher fidelity and higher resolution generations. I wouldn't be surprised if it takes the crown back from DeepFloyd when it releases, but I guess we'll have to see.
SDXL is available on StabilityAI’s hosted services already, so they can be compared head to head.
denoising in latent space certainly seems like the "correct" path. My (amateur) thinking is, the more you can do in latent space, the better.
The former seems likely to be lower compute-for-resolution, but that’s not the only consideration for “better”...
But there are consumer cards with 14GB+ VRAM, so its not, even before optimization, out of reach of consumer hardware.
>DeepFloyd IF works in pixel space. The diffusion is implemented on a pixel level, unlike latent diffusion models (like Stable Diffusion), where latent representations are used.
Automatic1111 is the defacto main UI for these kind of models. It will be supported there, quite quickly.
35% to full release by end of month, although it may not have adjusted.
I played with some ready prompts here
I found only this one on their subreddit: https://discord.gg/GvsvNrVkk5
But is harder to get a good picture. This fine tuned with a good RLHF will be amazing.
I've been trying to get a good portrait picture with "neon lights" on Stable Diffusion and it is almost impossible. Meanwhile with the new Dall-e, that was possible. The picture specially with SDXL is good, but it doesn't really have neon lights...
I tried now similar prompt on deepfloyd and managed to get there!
Explaining: I created a model of me, and wanted to create some good realistic portrait pictures. First I tried to create a model of me using some of the custom models already exist and the result was bad.
Then I tried SD 1.5/2.1... It was better, but couldn't really get some of the prompts make real...
Then I tried new Dall-e, saved, and inserted my face with img2img on SD and it worked much better!