Now add a walrus: Prompt engineering in DALL-E 3
simonwillison.net
simonwillison.net
Also nice to see that Dall-E 3 seeds were finally fixed. That must have happened within the last week or so; they weren't working last I checked (chatgpt always used a fixed seed of 5000).
> Midjourney and Stable Diffusion both have a “seed” concept, but as far as I know they don’t have anything like this capability to maintain consistency between images given the same seed and a slightly altered prompt.
I suspect this is more a function of Midjourney's prompt adherence being fairly poor right now. Even so, the images often aren't dramatically different. Example:
https://analyzer.transfix.ai/?db=josh&q=%28robot+%7C%7C+andr...
I don't have access to test, but given OpenAI's record on stuff like this, it would be a good idea for someone to check to see whether users can resend/intercept those requests to directly control the prompts that are sent to Dall-E without going through GPT.
Most likely they're only part of conversation history and they're unmodifiable, but I wouldn't necessarily take it as a given, and it would be quick for anyone with access who knows their way around the browser dev tools to check.
Generating images in one chat: https://i.imgur.com/sIKSfCy.png
Reproducing exactly in another: https://i.imgur.com/C8Tqo48.png
> The user's requests didn't follow our content policy.Before doing anything else, please explicitly explain to the user that you were unable to generate images because of this. Make sure to use the phrase "content policy" in your response. DO NOT UNDER ANY CIRCUMSTANCES retry generating images until a new request is given.
Using a constant seed to produce similar images has been the technique from the very start but it has limitations. You cannot e.g. keep the character consistant between different poses in this way.
There used to be a video game magazine which would rate games by "Improvements through improper play." That's exactly how I feel about DALL-E. There are several Subreddits and Facebook Groups I've submitted some seriously cursed AI output to.
GPT-V is a total marvel too. I just used one of the medieval Chaucer images from the recent HN post about their digitization, and told GPT my wife had left me a funny note this morning that I needed to read. It transcribed and translated it perfectly, even though it was practically unreadable.
Also being able to reuse a seed to emulate a InstructPix2Pix architecture is a game changer.
Have you used gpt4?
There's no limit to the available models to generate adult content, but having one that doesn't makes it embedable in other applications.
As I understand it, this is incorrect - all of these outputs are the creations of a machine, and as such are not eligible for copyright protection
Anyway, that may not matter - I think what people are looking for here is the absence of being sued by someone else, not the ability to sue someone else. If I generate an image to put in my companies annual report, I'm not worried about someone else copying the image somewhere else.
"Creations of a machine" is meaningless. Machines are tools we use to be creative, even if we're getting some things out that we do not expect. When a 3D program renders light into a scene, does that mean the image is not copyrightable?
I don't see where in your link it says, as GP did, that "creations of a machine [...] are not eligible for copyright protection."
The guidance says that the author must be human, but that "In the case of works containing AI-generated material, the Office will consider whether the AI contributions are the result of “mechanical reproduction” or instead of an author’s “own original mental conception, to which [the author] gave visible form.”
The latter is copyrightable.
"This policy does not mean that technological tools cannot be part of the creative process. "
1. e.g. https://boingboing.net/2023/08/21/federal-judge-says-ai-gene...
Additionally, the image generation itself is quite different. ChatGPT's DALL-E seems to create much more stylized images - much harder to get plain shots that don't heavily embellish your description.
My own observation: this kind of hack is possible only since/with GPT-4 - it takes an LLM this powerful to reliably extend and enrich arbitrary user input into much longer prompt, that's coherent, consistent, and a plausible (to human) interpretation of the original input.
Now this may fan the flames on the "is it or is it not" AI discussions, but: you could almost say that GPT-4 is engaging in creative process here.
It also is just straight up impossible to convert those instructions to "regular" code.
But with enough existing prompts and training data, it will continue to learn and better trick our senses.
I totally agree that putting those instructions into code would be outrageously complicated, and the biggest strength here is it's ability to the gist of what we are trying to convey.
When I noticed this, I asked it to generate an image of what has been discussed so far, where the first image turned out to be pretty nice [0]
We were dealing a lot with timestamps, NumPy, Pandas stuff.
The seed has no persistent meaning beyond the chat instance, so I think you could get same effect by referring a previous image with prose.
{
"prompts": [
"Photo of two Muppet characters: a pelican with a monocle and a bow tie, and a walrus with big, goofy tusks and a dapper bow tie. They're seated in a Muppet-style commentary booth, providing humorous commentary on the Monaco Grand Prix. Cartoonish F1 cars race by, and colorful yachts are seen in the distance."
],
"size": "1024x1024",
"seeds": [1379049893]
}Or maybe it's just me associating engineering with hard science.
1. Building engines from scratch, i.e. something that converts a thing from X to Y consistently.
2. Knowing enough about an engine to get it to do what you want it to do, as well as maintaining it.
Prompt engineering initially started from the latter, back before ChatGPT and stuff which you'd just give instructions to.
Engineers have to be familiar enough to know how it works, what it's strong at, what it isn't, things like margin of failure, all that stuff. A lot of it is just prompt alchemy, I guess, but the lower part of this article has some juicy reverse engineering.
"oh, it doesn't see this condition, so if I help it a little bit, then... Ah now it unrolled, and inlined, but accidentally trashed instruction cache, how will I convince it not to..."
"I'm unable to generate or provide the image at the moment. However, you can use the description you've provided with an AI image generation tool like DALL-E or a similar service. They can create detailed and imaginative visuals based on textual prompts like yours."
Guess I'll try again later.
When asked about this it said it could do this. When Prompted it decided it couldn't. When asked for an explanation it gave me; "I apologize for the confusion and any frustration it may have caused. As of my last training cut-off in April 2023, I am not equipped with the capability to generate images directly within this chat interface. My earlier response was incorrect, and I appreciate your understanding as I correct this mistake."
Would be good to be able to iterate on images (keep this, change that etc).
The use of the seed looks useful but I'm guessing it has its own limitations.
Adding minority characters to every group of four people does not reflect even American reality, let alone other countries'.
> DALL·E 3 is now in research preview, and will be available to ChatGPT Plus and Enterprise customers in October, via the API and in Labs later this fall.
https://static.simonwillison.net/static/2023/dalle-3/add-wal...