Imagen: An AI system that creates photorealistic images from input text
imagen.research.google
imagen.research.google
It seems to be the kind of bullshit statement that those companies put in place of "we paid $500k training this model and we're not giving it for free to anyone".
https://github.com/basujindal/stable-diffusion
You'd clone the fork, then download Stability AI's checkpoint, sd-v1-4.ckpt, from https://huggingface.co/CompVis/stable-diffusion-v-1-4-origin...
Follow the instructions in the forked repo, and you should be good to go in a manner of minutes.
I normally dev on Linux but don't do GPU work, only my Windows gaming PC has the horsepower to run stable-diffusion and unfortunately I couldn't get it working on Windows. Shame cause I'd love to play around with it! Does anyone have any tips?
Edit: Just noticed they removed the instructions from the readme for some reason. The requirements section can be found in this diff: https://github.com/basujindal/stable-diffusion/commit/487a0f...
Are you saying you coulda `conda create env -f environment.yaml` fine? This was my first conda install on Windows, so perhaps I missed something else.
conda env create --name envname -f environments.yml
Try reinstalling conda. Do you have a firewall on or anything like that?Maybe one of the newer commits broke something? I think I cloned it when b91816301fa62df239a45336b381ee918ff52b2a was the latest, but I'll have to check when I get home.
Maybe I’m being too optimistic but either way, given the pace of progress we won’t have to wait very long to play with this magic.
[1] https://openai.com/blog/reducing-bias-and-improving-safety-i...
It would be called out as fake or staged just like an imagen/walle2/openai would be called out as fake. The thing that makes stories real is real people - actual events and backed up testimonies.
As a side note, I'd watch that. I imagine the next scene shows a bloody room and a baby escaping.
> The thing that makes stories real
Since when do the hordes care about stories being real? People get harassed over photoshopped images. People get killed over false rumors. These tools make it easier to trigger such reactions.
They're literally just appending a race/gender string at the end [1]. In what world is that not just hot air?
[1] https://twitter.com/jd_pressman/status/1549523790060605440
Bias amplification is a real issue. https://www.theverge.com/2016/3/24/11297050/tay-microsoft-ch...
Tay might still be around if Microsoft gave a thought to potential issues before release. I’d prefer not to have this awesome technology tainted out of the gate as a tool for racists and pornographers. They’ll get their hands on it eventually but it’d be nice if they don’t get all the up-front press.
That is the only mitigation used in DALL-E 2, which up until recently was the only publicly available text to image model.
> I’d prefer not to have this awesome technology tainted out of the gate as a tool for racists and pornographers
Why is it your business what people do with the model? If people want to be racist they can already do so, they don't need a shitty model that doesn't work half as well as paying some guy in the third world $2/h to shitpost online. And I don't see the problem with pornography.
This section is just a requirement for some of the big ML conferences, like NeurIPS.
Google is in the business of ads, this might have useful applications in the ad space.
The only AI ethics I’m worried about are the lack of anything ethical going on w.r.t intellectual property in AI.
Scraping other people's content is already established fair use thanks to Google prevailing against the Author's Guild in front of SCOTUS. I find it difficult to understand how it could be illegal to scrape a bunch of creative works to create a system that generates new works, but legal to scrape a bunch of creative works to let people search two-page excerpts out of them.
Output copyright is very much up in the air. The Copyright Office has rejected copyright registrations claiming the software itself created the work; but presumably this wouldn't apply to a human taking ownership over something they used ML to create. The amount of prompt engineering you have to do to make these kinds of systems alone would count as some kind of creativity. The only real complaint I could see is if the system regurgitated its training data, which would be bog-standard copyright infringement.
Also, related note: I really hate how the whole AI thing is making the FSF sound like the Author's Guild did a decade and change ago. The law is already very clear that you cannot launder an infringement (copying GPL code) through a fair use (trained ML weights). Please do not adopt the arguments of copyright maximalists.
This is just a link to their previous posted site btw, nothing new.
Using ML to determine the order of matches is absolutely fine, but to "digest" the internet and cook up an answer to what I'm looking for without proper sourcing? I do not want that. I don't want to try and guess what biases the language model might have. It's way easier for me to gauge the bias of another human, and for that, I need to be sent to a page a human has written.
(Of course, I realize that the "blog written by a bot" genre of writing is also becoming more convincing, making this whole thing harder...)
I already feel like google search sucks. It is partially googles fault partially the internets fault. The best source of information was and will always be forums full of knowledgable people having genuine conversations and everybody moved to discord because.... because... om I dont know and besides google doesnt get paid to elevate tiny ass forum websites in the search results.
It’s like the “AI” projects that spend all their time building a fancy sci-fi-looking robot body, ignoring the fact that how much of a “person” a robot seems has almost nothing to do with how physically anthropomorphic it is. Johnny 5 and Wall-e are proof enough that the mind, not the geometry of the body, is what’s important.
AI content generation (text, image, source code, video, music) will be a huge boon for prototyping where applied judiciously.
is it that good now? note that Pixlr was bought by Google
Did that happen recently? It says in Pixlr's site they're owned by INMAGINE.
Google's product is vaporware and we shouldn't afford them any airtime until they release something usable. They're just trying to butt in and get press off of the backs of the teams actually working in the open, and that's super lame.
Release your model, Google, or stop bragging and talking over the others here. You're greedily sucking oxygen out of the conversation, and as a trillion dollar monopoly you don't deserve anything for free off of the backs of others. Not when you're not contributing. Stop being the rich kid talking over everyone else about how awesome your toys are.
Anyhow, the real story is Stable Diffusion. They're actively demonstrating the correct way to run this as opposed to the entirely closed OpenAI DALL-E or the (again vaporware) Google non-product.
Even MidJourney uses Stable Diffusion under the hood, using sophisticated prompt engineering to make their product distinct and powerful.
There are very good reasons for withholding SOTA models, primarily from the info hazard angle and avoiding escalating the capabilities race which is basically the biggest risk we have right now.
Google / Deepmind have actually made some good decisions to try and slow down the race (such as waiting to publish).
Take a look into scaling laws and alignment concerns, this is a very real challenge and existential risk not some crackpot theory.
What good does a few months lag do when nobody is bracing for impact?
Even ignoring the infohazard angle if they published everything immediately that would escalate the race. By sitting on their capabilities and waiting for others to publish (e.g. PaLM, Imagen vs GPT-3, DALL-E) they are at least only playing catch up.
Whatever effort Google has put into building the model is infinitesimally small compared to the work of the creators they're harvesting.
I don't expect this to happen easily, if at all, but I'm strongly in favor of it, and would even support legislation to that effect.
It worked okay - one issue was DALL-E wants to keep "style" consistent so any stray bit of debris greatly affected the interpretation, but I did in fact get 1 design idea out of it which changed how I think we'll do a bit of it.
These things in many ways are just extremely enhanced search tools - "describe what you want to see"
You have midjourney.com beta.dreamstudio.ai craiyon.com (real quick version no fuss, low quality) creator.nightcafe.studio
and those are just some of the entry-level ones. I make ai art all day long! Check out my media feed on twitter @Sheilaaliens
I feel like this is the epitome of modern "content creation"
Typing a few sentences into a software you barely understand, said software shits 15 jpegs out, 2 are good, "hey I make art". What a sad state of affair, tech is consuming everything and people are cheering, one more step on the path to being complete useless key pressers.
On the other hand, the whole point of automating creation is that you don’t need to understand the underlying mechanisms.
https://medium.com/org-hacking/pioneers-settlers-town-planne...
When your tool becomes a megacorp owned subscription based service you barely understand is it really a tool ? It doesn't produce art in the way a brush and paint produce art, you barely have any control on what it does, you just become an image filter, you press a few keys, look at the image for 5 seconds and decide if it triggers the right part of your brain or not
It's not so different from the artist shitting on canvas, but at least he had the creativity to do something new and daring.
The end result might be called art but the whole process is completely devoid of what makes art "art"
Owning the tool vs being owned by the tool, yadda yadda..
You have quite a lot of control from style and colour to composition by feeding in initial images. You can do this iteratively, selecting parts you want to keep and others you want to adjust.
Photography hasn't killed art, yet you can describe it as "point at something and press a button". Sure you'll get something out of it, and it might be alright. But the great outputs take more work and care, just like with the air art now.
Press a button on a box you barely understand and the camera shits out a pic. "Hey I make art". What a sad state of affair, tech is consuming everything and people are cheering, one more step on the path to being complete useless key pressers.
Also if you are doing something interesting with digital photography it is definitely not just pressing a button.
At worst the complaint seems to be "with care you can more easily create good art" which is a very odd complaint.
You can tell that right away seeing how many camera users never produce anything of quality. A 6 years old can press the button of a camera but no 6 years old will produce meaningful work. In the case of Ai art tools you just need some basic english skills
You still have to select the 2 AI jpegs out of the 15, that's the artistic part.
> it's still incredibly more complex than "hey google draw me a horse with a tuxedo"
And just like taking a photo of a random horse in a field you'll mostly get a pretty bland result.
You can argue it's simpler to create art with, which is an odd complaint, but if you're not working to make something great you generally won't get it - just like photography is much simpler than painting as it's "just press a button".
I think this is an oversimplification.
There was a recent write up by a guy who used DALL-E to create his logo for his open source project. What was clear from that writeup is that it is still a process for getting exactly the look that one is aiming to achieve. Even with an AI, there are different styles, decisions, choices, and visual representations that have to be made.
Your position that with photography, you have to "chose a subject, lights, &c" doesn't change with AI generated artwork; one still has to describe the subject, color scheme, visual style, composition details, etc. for the AI to generate the image. Except that instead of composing a scene with makeup, props, and subjects, you do it textually.
I'd say that in some respect, it is far more "creative" than photography because it removes physical and real-world constraints from the artist which would otherwise require knowledge of CGI and digital tools.
> A 6 years old can press the button of a camera but no 6 years old will produce meaningful work
This is also true of even painting. Even a 6 year old can grab a paintbrush and paint without producing meaningful work. So that does not change with AI generated artwork. Yes, a 6 year old can describe a scene to an AI that generates some image -- just as a 6 year old can pick up a brush and apply paint to a canvas, but the likeliness of a 6 year old presenting the seed/input that the AI needs to generate something unique and of visual interest/originality is low just as it is with a paintbrush.
One of the truly fascinating things with stable diffusion is that you can use a starting image. So you can start with a vague sketch to control the composition. It's quite incredible.
Not really. This very much depends on the prompt. Just look at these, made these yesterday in a batch run. Same prompts, different seeds. This all what it made on seed it chose.
The keypressing is clearly, objectively, not useless. It is useful, just less costly and more accessible to way more people. I'm generally happy with this.
No models to pay and to check for legal ages, infinite possibilities: just write the pic you want, the massive body part you want, how many genitals are involved and boom. You have your image.
[0] (OBVIOUSLY NSFW) https://news.ycombinator.com/item?id=32572770
10 min * 60 * 24 FPS = 14400 images ~ 2^13
So maybe 10 years?
can someone in the field explain why this has exploded recently? there seems to be a lot of these tools released recently (text to image) was there a major breakthrough? a new idea that pushed everyone forward? a recent sharing of talent between groups?
edit: just another thought, are they just being posted to HN now, i don't see a date on the page for when it was released . I also don't know the general term to find a list of all of these to find all the release dates
Here are some of the projects on GitHub: https://github.com/topics/text-to-image
Another good source is https://paperswithcode.com/task/text-to-image-generation
The moment everyone knew this was going to be big was in 2019 when StyleGAN came out. They used a lot of tricks like aligning face features (like eyes) and had all their pictures of a single domain (the most famous being faces) but none the less, that was the moment everyone in the AI field knew this was going to be big, and so three years ago a lot of big people shifted to this line of research.
The four main innovations since then have been:
1. Transformers
Generalized computation kernels which allow for images to consider non-localised relationships between pixels of an image. Released in 2017, and originally used for language.
2. Pixel Patch Encodings
Different resolution semantic and geometric image information encodings which allow for better representations of relationships between image areas than pixels are able to achieve given the same compute. Allows using Transformers on high resolution images.
3. CLIP
Contrastive Language and Image Pairing. Before, the only way we knew to classify an image was as a "face" or "cat" or "ramen". When the genius idea of labeling images as semantically meaningful vectors rather than one hot encoded classes was revealed, it changed everything in computer vision very quickly, and problems that used to be hard became trivial. Released in 2021
4. Diffusion Models
GANs penalise you for making an image which does not seem to be part of an existing dataset. This encourages one to make the worst quality image that looks like a member of that dataset. Diffusion learns to denoise an image, removing noise is perceptually similar to increasing resolution, people like images that look that way. There may be more people with better intuition about diffusion models may be able to add on why they're superior. I've read all the papers leading up to the latest unCLIP (Dalle2) but it's complicated. Released in 2020, with major improvements to the training process continuously being made since then.
Hope this was helpful. All of the above were only implemented for images in any real way in the last three years. Putting them all together is something many people only just this year did, resulting in DallE, Stable Diffusion, and Imagen.
I'm working on doing this for 3D and later for use cases in AR. 3D generation still hasn't been cracked the same way image has but the above will likely contribute to the solution to that as well. Anyone who's intersted in working on that feel free to message me.
P.S. It seems raccoons are unimaginable (even for AI) with any sunglasses: if photo-realistic mode is selected for a raccoon, changing to "wearing a sunglasses and" makes no difference :)
The models are a product of their datasets, specifically the relationship of the images and prompts via CLIP. CLIP puts both images and text into coordinate space, imagine just a 2D graph. It tries to assure that for any real image and its caption, they will each be each others closest neighbor in that coordinate space.
So if you want a certain image, you have to ask "what caption would be most likely and most uniquely given to the image I'm imagining".
I'm sure this advice is way less helpful than what you find in prompt engineering discord channels and guides I've seen.
Secondly, there's vastly more labeled image data in the world than 3D data, so creating a CLMP (contrastive language and mesh pairing) model is harder.
It's very late but I may be able to give a much better answer on more of the nuances of 3D generation tomorrow.
I can imagine you'd have the problem of stray floating voxels then, which isn't as noticeable when it happens with 2D pixels.
I knew about transformers, CLIP and diffusion, but pixel patch encodings are new to me.
Can you give me more details / point me towards an explainer? A quick duckduckgo search didn't help.
What does that mean?
(Thanks for the explanation)
The models behind Imagen and StableDiffusion are actually simpler than DALLE2, and both are higher quality (SD of course isn’t always since it’s much smaller). That suggests DALLE3 will also be simpler again.
There’s also been very recent work with generalized diffusion models (that use problems other than noise removal and still work) and Google researchers have been tweeting results from a merged Imagen/Parti in the last few days.
For example, here is an RPG designer using Midjourney for illustrations: https://www.bastionland.com/2022/07/primeval-bastionland-pla...
One major thing that happened recently (2ish weeks ago) was the release of an algorithm (with weights) called stable diffusion, which runs on consumer grade hardware and requires about 8GB of GPU RAM to generate something that looks cool. This has opened up usage of these models for a lot of people.
example outputs with prompts for the curious: https://lexica.art/
There is a huggingface instance, Collab notebooks, and local running notebooks here. [1] on the stable diffusion subreddit.
Also someone has packaged an exe that runs it with no fuss on computers with Nvidia GPUs that they posted on the media synthesis subreddit[0]
In my limited testing this compares ok to Dalle2. Style shifting works slightly less well and it's hard to force it away from normal images but with a little work it tends to be more accurate to your prompt.
[0] https://grisk.itch.io/stable-diffusion-gui
[1] https://www.reddit.com/r/StableDiffusion/comments/wqaizj/lis...
I think "what's next" is fitting these tools together in a larger system
* Text to 3d model
* Text to video clip
* Illustrations for newspapers
* Generating pictures for food menus
* Generating a music video from a song
* Generating pictures or even an entire movie for a book
* Interior design ideas
* Product design ideas
I have noticed that leveraging these tools isn’t easy. It requires a fair bit of creativity to come up with the prompts to create an image to really wow somebody.
One murky area we're still far away from but I'm curious to follow the developments on: AI-generated movies. Once generated clips gets good, what if some movie buff can just generate a movie, scene-by-scene, just using these tools? What about "casting" certain celebrities? The comic I mentioned uses Zendaya (probably because she plays a character in the Dune movies) as a character
Here's a slew of images (1 through 5) I generated all from the same seed and same prompt sans a word or two: https://www.instagram.com/p/Chg60Fou6xB/
This is what I've been doing for any recipes[0] that don't have pictures with pretty good result.
[0] www.reciped.io ex: https://www.reciped.io/recipes/mushroom-and-onion-pizza/
Smart for Google to invest in this because their business relies on third party content to exist (blogs, webpages, youtube videos)
If they can vertically integrate their business to create the content AND own the discovery algorithms then they officially win the internet
FWIW, I don’t think the AI systems will generate a whole video by itself - it’ll be some form of image to image generation where an artist will render a rough sketch of the scene and the AI will fill in the details, frame by frame.
In general, I think these models are a great and funny toy, but not a threat to stock-photos yet. This may change within a year or three years though.
crystal dragon thing:
https://cdn.discordapp.com/attachments/951197655021797436/10...
https://cdn.discordapp.com/attachments/951197655021797436/10...
https://cdn.discordapp.com/attachments/951197655021797436/10...
davinci-style notebook of flying machines:
https://cdn.discordapp.com/attachments/1008049109338443829/1...
https://cdn.discordapp.com/attachments/1008049109338443829/1...
this person tried to show the life cycle of an alien:
https://cdn.discordapp.com/attachments/1010211132671275058/1...
https://cdn.discordapp.com/attachments/1010211132671275058/1...
https://cdn.discordapp.com/attachments/1010211132671275058/1...
https://cdn.discordapp.com/attachments/1010211132671275058/1...
cavemen taking a group selfie (lots of faces)
https://cdn.discordapp.com/attachments/1011408429170044928/1...
Also I realize there's a lot of image prep work required, not to mention I have a less than ideal amount of VRAM (3060 Ti w/8GB but no monitors attached i.e. 8GB free) so I have to lower some settings. The source images have to be in 1:1 format (which none of my photos are) so I'm using a script to batch call ImageMagick's 'convert' to add white borders to the top/bottom, which results in my renders also having white borders.
[0] https://docs.google.com/document/d/1CnC5SaqpeJiQS-TlDS4trzJR...
Here's a couple "first results" that I personally tried and you can judge for yourself:
"the government is putting violence in our water"
https://mj-gallery.com/87f5a54d-7d59-44d3-aab4-1dd3ef34902e/...
"cherry monkey"
https://mj-gallery.com/5d2e14ba-8ea1-4797-ab6b-4a6807cfffa8/...
"permaculture garden city"
https://mj-gallery.com/088c18c1-8e61-44da-b109-edfbd32967ac/...
"Acmella oleracea"
https://mj-gallery.com/9d1bc9f3-3cdc-44ec-8577-0791c69aa942/...
"mondrian banana cloud forest"
https://mj-gallery.com/25a8ef06-1a07-4a40-97e5-bebbd2ee925e/...
"lonely neon rainforest at night"
https://mj-gallery.com/8e7b73c0-f519-4727-bf16-30e0a42ab412/...
Obviously these prompts are a bit more artsy than stock photos are meant to be but the point is just to give you an idea of how it does on the first try. All of these took less than a minute to produce
[0] pornpen.ai: https://news.ycombinator.com/item?id=32572770
It’s easy to remove the filter from the SD scripts and that’s intentional.
https://media.discordapp.net/attachments/999426920376717513/...
Generated with Midjourney Beta
As far as I could tell (using it before, during, and after this Beta option was available) it was the upscalar using that, not the original 4 image generation.
Our brains are very thoroughly wired to detect faces in general, and flaws in faces. So we have a very, very high standard for what passes muster.
We are apparently much more forgiving with regards to what raccoons and corgis look like.
To learn logical concepts just from images seems entirely impractical, like we can't rely on having enough images such that models can understand words coherently as language. You could draw a picture of a sign that says "children crossing" not because you can understand and remember exactly what an image of such a sign would look like, but because you have an understand of English and the character set that would let you reproduce it. If you tried to learn to create the same sign in Arabic you'd either need to see a huge number of signs to learn from or (more likely) build a language model for Arabic.
The kind of abstract understandings that we know we can train in language models just aren't learned by image transformers at this scale (or likely any practical scale). A language model could easily understand: "A red cube is stacked on top of a blue plate, a green pyramid is balanced on the red cube" and infer things like the position of the pyramid relative to the blue plate, image models quickly fall over with such examples.
An interesting nascent (and hacky) example of the benefits of combining models is people are using language models like GPT-3 to create better prompts for image models.
I think a lot of this has been solved in DALL-E already. It's pretty good at right-looking faces, and fingers. Text not so much... but it does appear that whatever OpenAI are doing behind the scenes, that's getting better too.
Also, I seem to recall that at least some models deliberately harmed generation of human faces (e.g. by selection of training data) to draw away attention from the deepfake/fakenews usecases and the related ethical,political and PR issues; I would assume that if any of them wanted to actually try and make specifically faces look good, that would be purely a matter of some engineering work without any breakthroughs needed - I mean, we have evidence from face-specific models that the same technical architecture can do decent faces.
https://replicate.com/laion-ai/erlich
It still has issues of course, but a lot better at spelling than DALLE2.
The major breakthrough that happened 6 months ago is that someone put their api behind a website for people to play with
The porn star one said it had used up its resource, and this one I cant figure out if i can run it. I get spreadsheet from the linkm but I am unsure how it helps me.
As a guy who does not know a lot of "AI" are these things just showing they "Hey I trained this then to do ...... / I have a big data model"?
Is there a standard way to share it?
If you start out with 10.000 penguins to teach a model about penguins. Once the model has worked through it, do you have any more use of the images separately?
How large are the models?
In terms of theory, such systems are candidates for a generic perception engine you might use in say, a robot with cameras, speaker and a microphone.
Perception is just one aspect of intelligence, but this research ultimately makes it possible for a machine to encode data semantically.
2022: I made a new AI image generator!
It's another google project using a different set of techniques.
Parti+Imagen is in development and is competitive again.
Ex. Show me walking on a beach. Show my dog wearing sunglasses, etc.
https://textual-inversion.github.io/
It allows you to find a string that corresponds to some "thing" you have pictures of. Then you can use a text to image model containing a reference to your "thing".
Looks pretty amazing! However it needs 20gb of VRAM so I couldn't try it out.
I wonder what's next?
But seriously it may usher in a new cottage industry of content creation, from graphic novels,to animation to movies.
Stable diffusion was un-neutered within 24 hours of its public release and the worst people do with it is Emma Watson porn.
A more cynical side of me just thinks Google is rushing out the PR (including the hand-plucked sample images) because they can see that the state of the tech is progressing rapidly and perhaps by the time their tech is release-ready a competitor will already have something better (the new round of betas certainly look promising.)
It is a little bit on brand for Google to make an announcement they have the best, only for those claims to fall over later.
The technology isn't special to Google. They don't control its proliferation.
EDIT: Just found this[1] as well, though setup might also be a pain.
[0] https://softology.pro/tutorials/tensorflow/tensorflow.htm [1] https://replicate.com/methexis-inc/img2prompt
Microsoft has an app called Seeing Eye that’s not as good at it.
Both of those use older pre-CLIP technology.
Filing a charge is pointless, says Ezra. Since two years, she's being harassed on Telegram. It started when she was sixteen: photoshopped nudes with her snapchat account were circulated. They had taken selfies from her social media, and those of her family, and combined them with porn fragments. She doesn't know the perpetrator, but that person takes a lot of trouble to ruin her. "Nowadays, the boys have so many ways to make it look real." [1]
If you read that, and think all these tools should be released, you're part of the problem.
[1] de Groene Amsterdammer,146/33, p. 21.
Such comments make me wonder whether a social credit system is actually a good idea. Yes, it can be abused to deny my rights, but how can it be worse than being assumed to be a sexual predator by default?
But your comment makes it sound as if you'd rather give up your rights than not have access to this system. I don't think it's that interesting, is it?