DALL·E 2 prompt book [pdf]
dallery.gallery
dallery.gallery
Someone else here on HN observed that everyday people don't "get" how huge this all is. I experimented with asking random acquaintances at a local cafe for prompts and showed them the generated pictures. All but one person was totally unimpressed.
If everything feels like magic, then what's one more piece of magic?
This scared me more than the implications of DALL•E 2 itself specifically. People think of the technology in the world as mysterious black boxes that do inexplicable things, and they hence no longer understand relative complexity, progress, or change.
My impression is that to most people DALL•E 2 is not "substantially" different to, say, Google Image search. Text in... image out. What's the big deal?
Edit: ah, I think I was thrown off by the use of “median/average” as equivalents.
How do you know something is impressive or surprising? You compare it to the previous status of the industry, which is something you know, but the random people in the coffee house don't.
I doubt they really believe it's truly inexplicable. Most layman know that somebody knows what's going on in their device.
And nobody has an explanation for everything they use, but it doesn't mean they get attributed to magic. I have no idea how a bridge is designed and built, but I don't believe it's magic, it's just something beyond my knowledge.
I agree that People think of technology as black boxes that do inexplicable things, but I think that only matters if boxes interferes with their contexts and scopes. It seems it is understood a threat to digital artists, but at the same time, it is possibly less relevant than Google Search as it is.
.-- .... .- - / .... .- ... / --. --- -.. / .-- .-. --- ..- --. .... - ..--..
(WHAT HAS GOD WROUGHT?)
https://code.flickr.net/2014/10/20/introducing-flickr-park-o...
Today a hobby programmer could do it by themself.
I think this is because people reason this way: the computer has many pictures in its archive, saved as "dog.jpg" and "hat.jpg". If your prompt is "dog with hat", the computer just combines them. I think this is what people think is happening, so they are less impressed.
Over the past 7 days I’ve generated ~1000 images, 150 of which were good enough to save. I only saved images which made me audibly gasp.
Witnessing your own novel idea spring to life is a magical experience. DALL•E provides an artistic tool on a comparable level to digital photography, and by extension Photoshop.
At this stage it’s 100% clear to me that DALL•E has heralded in a revolutionary new age of design. Every day I worked with it, I grew more confident in my outlook.
It might not necessarily be an OpenAI product which truly “integrates” with humanity — but DALL•E has shown me that it’s possible… and just a matter of time.
I'd happily pay for it, but have a hard time figuring out how.
The relevant code is linked below and is a mess, but the idea is: 1. Generate a base image 2. Use inpainting to expand the left-hand side of the root image. You can do this by submitting an image whose left side is transparent and whose right side is the left side of the root image 3. Ditto for right-hand side 4. Stitch the three separate images together https://github.com/charlesjlee/twitter_dalle2_bot/blob/main/...
It's not as good for doing concrete asks, but it's very good for getting specific vibes.
The website feed [1] requires Discord login to view examples, but there's some unofficial galleries [2]
[0] https://www.midjourney.com/
It's better in some ways than craiyon and worse in others from my sampling (craiyon seems to give me less 'nice' options but better logo-y images)
* $10 for ~200 prompts * $30 for ~900 prompts + unlimited "slow" prompts where your job is put at the end of the queue and you have to wait longer (no idea how much longer though.. are we talking about seconds or hours here?)
I just got access to Dall-E today, and there it's 115 prompts for 15$, so roughly twice the price.
It's much like hiring a painter, a writer or a web designer... .
At least, it's interesting for me to see what may be the beginnings of a new job class. I've been wondering where all this ML/AI business might take us.
Some similar examples below but you will need to engineer the prompt a bit more to get it exactly the same.
First off, the PDF in this very post, as well as promptwiki link posted elsewhere in the comments, shows that there are definitely people doing this for free too.
As for your analogy, I'd say it's closer to paying someone to help you find something obscure and hard to find on Google. Anything that requires skill and time to do should be allowed to be monetized. Of course, as you mention, there will always be people doing and posting it for free, but I don't see why people shouldn't be able to make money from something they've taken time to master.
Dall-E is a tool, and this is like hiring an expert that can use that tool effectively. It's no different than hiring someone to Photoshop something for you, or more precisely here, paying to download/license premade Photoshop content someone put time and effort creating.
That's exactly what google dorks were. Google started out as a great search engine, but it's also a tool and 'Google Fu' was a real thing. Anyone could easily find song lyrics and news articles. It took effort (time and skill) to understand how google worked and how to format searches to get the results you were interested in for many obscure types of data. I'm not suggesting that people shouldn't have been allowed to charge for Google search terms, only that unlike today making a fast dollar wasn't anyone's first priority.
A marketplace for Google search terms might have been just as successful, but it never happened and I'd like to think we were all better off for it. Instead that information was shared freely and widely to anyone interested and it served us well for many years until Google degraded their product by no longer following their own rules and much of that became useless.
I'd argue that this is much less like paying someone to create something in photoshop (Here you give the right commands, DALL·E 2 does all the work) and much more like charging for a tutorial on how to use photoshop to achieve certain effects (You follow the right steps, photoshop does all the work). There's nothing wrong with charging for tutorials (many do and have done), and I'm glad that there are still people willing to share what they've learned without throwing up a paywall, but the rush to monetize here was shockingly fast. An entire infrastructure was put into place to support buyers, sellers, purchases, featured products, and payments before many even had a chance to try this new tool for themselves.
I see it as a reflection of how much the culture of the internet has changed. The commercialization of the web has become so pervasive that many people can't imagine an internet without a profit incentive (often expressed as some form of 'the internet couldn't exist without ads!' or 'No one would create content if they aren't getting paid to do it!'), but those of us who are older will remember a healthy and thriving internet that existed long before it was commercialized and how many useful, popular, and helpful websites and services were created and maintained without any thought given to "Yeah, but what's in it for me?" and the overall vibe on the internet was about sharing (or more cynically, showing off) vs making money.
I still maintain that Dall-E is just a tool just like Photoshop, it just happens to be a layer above. IT takes time and effort to find the right prompt to get the result you want, just like it takes time and effort to find the right effects and filters in Photoshop to get what you want.
Anything that takes time and effort to do will always have demand, which will always have people trying to monetize, while others trying to do for free.
In my mind, the main eras of content on the internet look something like this:
Epoch 1: Pure, unblemished user generated content. Message boards and forums rule.
Epoch 2: More user generated content + a healthy mix of recycled user generated content. e.g. Reddit.
Epoch 3 (Now): Fake user generated content (limits to how much because humans still have to generate it). e.g. Amazon reviews, Cambridge Analytica.
Epoch 4: Advanced generative models means (essentially) zero friction for creating picture and text content. GPT3, Dalle-2.
Epoch 5: Generative models for videos, game over.
IMO, the future of the internet feels like a totally disastrous (un)reality. If addictive content recommended by the likes of TikTok has proven anything, it's that users ultimately don't care _what_ the content is, as long as it keeps their attention. It doesn't matter if it comes from a human or a machine. The difference is that in a world where the marginal cost of generating content is essentially zero, that content can and will be created and manipulated by large malicious actors to sway public opinion.
The Dead Internet Theory will fast become reality. This terrifies me.
[1] https://www.theatlantic.com/technology/archive/2021/08/dead-...
Healthy communities support artists. Generative models aren't truly creative art, they are explicitly and exclusively derivative. They are exactly inspired, in a different sense than artists who are inspired.
Artists can use these new tools, and so can non-artists, and even if the resultant image is the same, I think there is a difference depending on the intention of the prompt-er.
These models are democratising content creation, not art.
I dunno. Weird and very scary.
I spent a lot of time alone on airplanes when I was a young father and there’s something bittersweet about the solitude and beauty in this image for me. My favorite parts about this image are the gradient in the sky, the waning sunlight in the top corner and the very faintly illuminated frame around the entire window.
Very happy with the print. Next time I might get the satin finish though, it’s like a mirror.
That's why it's really nothing but a PR exercise. I honestly don't think they care much past that.
AIs are going to replace entry level creatives, and experienced users with taste will largely perform selection and the development of good starts to mature designs. And I mean all creatives. Engineers, architects, mathematicians, programmers.
Programmers, also doubtful. You still need to design some APIs and interact with them. Interacting with an AI using natural language might become possible, but it definitely won’t be as efficient as more structured languages. (E.g., writing an algorithm in actual code is often much easier than teaching it to a human.)
Someone had taken a photo, somehow digitized it, distributed it, and we were looking at a representation good enough that we could tell what it was.
It felt like we were living in the future - me as a middle schooler and him with decades of software development under his belt.
The iPhone maps app with the GPS dot and DALL-E are the only things that have matched that feeling.
I come from city that used to be known for cheating taxis… at that point I knew I will never go back.
My favorite is coming up with two word prompts, like "endless beginnings" or "stressful shapes" or "happy anxiety".
When I got access on Sunday, I first tried a lot of different prompts and got some interesting results. One semirandom one, “A photograph of a professor playing a grand piano on a rainy night in Tokyo,” produced some very atmospheric images. I then went down a rabbit hole of variations on that prompt (“A painting of...,” “A line drawing of...,” “A painting in the style of Rembrandt of...,” etc.).
I put most of the results into the following video, if anyone is interested.
Please feel free to take screenshots of the images in the video and share them yourself on the Internet—giving appropriate credit to OpenAI, of course. You don’t need to credit me.
You may also run the same prompts through DALL-E 2 yourself and post the resulting images online; those images will be different from what I got. Or you could come up with your own prompts and post those images. In either case, please post the link here. I will be very interested in seeing the images you get.
I liked the music! Didn't know it was yours :)
I'd still prefer to see the images in a way that enables me to choose how long to look at each. It could be a webpage that plays your music in the background?
> Please feel free to take screenshots [...] You may also run [...]
Well, I wanted a low effort way to view the images, but thank you :)
This twitter thread[1] also has some good suggestions and an interesting approach.
[1] https://mobile.twitter.com/fabianstelzer/status/155422934750...
NPCs that sit in your room and provide individual-specific training on niche topics. Provide talk therapy to overcome issues, or act as an assistant in helping you explore, research and document new fields.
Play a part in an episode of a vintage sitcom, taking it in an entirely new direction. View the rest of the season based on the changes you've made.
Progressed AI tools combined with improved human-computer interfaces will introduce amazing possibilities.
If you post those images online, they seem to ban you.
>Use of Images. Subject to your compliance with these terms and our Content Policy, you may use Generations for any legal purpose, including for commercial use. This means you may sell your rights to the Generations you create, incorporate them into works such as books, websites, and presentations, and otherwise commercialize them.
I assumed that means you can derive on the generations as well. E.g. when you are creating game assets for yourself, you won't want to have the watermark on them in the game or screenshots of the game (which may be published on the web).
---
Also, how can it even work? I take a Generation that you posted on Instagram, crop the watermark, reupload it and you will get banned?
Some will say that this 'tool' will help you in your creative process. In my obviously "biased" view, in the next 2-3 years, this will lower your monetary reward in half and create more requirements for competition with AI. People already are comparing DALL-E with human results.
In the long run, the software industry will eat itself to oblivion. Greed has no boundaries, and optimization of costs for corporations will never stop.
Some of my art related colleagues saw this 'trend' early and pivoted to crafts with added value for customers in the real world. On the oil painting side (which I am at) I don't feel any form of pressure, I paint for myself as a therapy. So, good luck:)
Whether or not this particular iteration of the model is 'good enough' to be widely applicable, or whether DALL-E 2 is 'creative', it's only a matter of time before the way humans interact with media is changed profoundly.
I always thought CRUD work might be automated eventually, while I would have guessed that something like embedded/high-performance was pretty safe, but now I'm not so sure anymore...
The images give a good first impression. Which is... impressive in itself. But they won't fool anyone who's studying them for even a few seconds.
I love Dall-E but I still use other models. Some of my favourite results have come from JAX CLIP Guided Diffusion:
https://colab.research.google.com/drive/12Bod44YVIXYRh39WRqp...
Disco Diffusion still holds up for painterly stuff. Majesty is great for portaits. MidJourney can beat/match Dall-E for lots of styles. I got very good results from Multi-Perceptor VQGAN+CLIP v4 for matching artist styles.
etc etc.
Dall-E is amazing and versatile but it's often lacking some "soul" that I get from other models.
I am signed up for the waitlist and can't wait to give them my money.
I JUST got my invite and was googling prompt suggestions. The timing on this article is incredible.
and if any tech layer fails, it all fails
example - battery shortages, due to labor shortages, due to covid, due to...
I joined a month ago with no invite yet so it may take some time.
AFAIK it is free either after you get in. You will need to buy some tokens to generate images. See: https://mixed-news.com/en/openai-announces-pricing-for-dall-...
It’s a cool thing that creates cool things. How does “maybe something else was their first” effect that at all?
DALL-E had a million things that had to be there first before it could do its magic. How does that fact take away from the excitement of the new realm of capability we have access to today?
I think it’ll happen eventually.