Spent $15 in DALL·E 2 credits creating this AI image
pub.towardsai.net
pub.towardsai.net
I had the same trouble. In my experiment I wanted to generate a Porco Rosso style seaplane. illustration. Sadly none of the generated pictured had the whole of the airplane in them. The wingtips or the tail always got left off.
I found this method to be a reliable workaround: I have downloaded the image I liked the most. Used an image editing software to extend the image in the direction I wanted it to be extended and filled the new area with a solid colour. Cropped a 1024x1024 size rectangle such that it had about 40% generated image, and 60% solid colour. Uploaded the new image and asked DALL-E to infill the solid area while leaving the previously generated area unchanged. Selected from the generated extensions the one I liked the best, downloaded it and merged it with the rest of the picture. Repeated the process as required.
You need a generous amount of overlap so the network can figure out which parts is already there and how best to fit the rest. It's a good idea to look at the image segment you need to be infilled. If you as a human can't figure out what it is you are seeing, then the machine won't be able to figure it out either. It will generate something, but it will look out of context once merged.
The other trick I found: I wanted to make my picture a canvas print, and thus I needed a higher resolution image. Higher even then what I can reasonably hope with the above extension trick. What I did is that I have upscaled the image (used bigjpg.com, but there might be better solutions out there.) After that I had a big image, but of course there weren't many small scale details now on it. So I have sliced it up to 1024x1024 rectangles, uploaded the rectangles to DALL-E and asked it to keep the borders intact but redraw the interior of them. This second trick worked particularly well on an area of the picture which shown a city under the airplane. It has added nice small details like windows and doors and roofs with texture without disturbing the overall composition.
What I did:
With the prefix of the prompt I described the image. I started the extension operations with "red seaplane over fantasy mediterranean city" but then I quickly realised that this was making the network generate floating cities in the sky for me. :D So then I varied the prompt. "red seaplane on blue sky" in the upper regions and "fantasy mediterranean city" in the lower ones.
I went even more specific and used "mediterranean sea port, stone bridge with arches" prefix for a particular detail where I wanted to retain the bridge (which I liked) but improve on the arches. (which looked quite dingy)
(I have just counted and it seems I have used 27 generations for this one project.)
Maybe Dalle-2 is just secretly a studio Ghibli/Miyazaki movie fan.
I was testing to see how close I could get to replicating a t-shirt graphic concept I saw.
I had been using ~"A telephoto shot of A neglected police car from the 1980s Viewed from a 3/4 angle sits in the distance. The entire vehicle is visible but it is overgrown with grass and flowery vines"
This process sounds great, though it seems like DALLE needs to offer tools to do this automagically.
Here is "llama in a jersey dunking a basketball like Michael Jordan, shot from below, tilted frame, 35°, Dutch angle, extreme long shot, high detail, dramatic backlighting, epic, digital art": https://imgur.com/a/7LoAtRx
Here is "Llama in a jersey dunking a basketball like Michael Jordan, screenshots from the Miyazaki anime movie", much worst: https://imgur.com/a/g99G7Bn
https://parti.research.google/ https://imagen.research.google/
The models themselves are not public however.
Images are more artistic and less clip art-like than Dall-E, but also don’t have a house style like Midjourney. It’s stunningly good - and open source.
What’s really cool is that the devs have worked hard to optimise the model, so after being trained on 1000 A100s it’ll run happily on an 8gb graphics card or M2 Mac.
But this DALL-E thing, it's really blowing my mind. That and deep fakes, now that's sci-fi tech. It's both exciting and a bit scary.
The idea that in the not so far future one will be able to create images (and I presume later, audio and video) of basically anything with just a simple text prompt is rife with potential (both good and bad). It's going to change the way we look at art, it's also going to give incredibly powerful creative tools to the masses.
For me the endgame would be an AI sufficiently advanced that one could prompt "make an episode of Seinfeld that centers around deep fakes" and you'd get an episode virtually indistinguishable from a real one. Home-made, tailor-made entertainment. Terrifyingly amazing. See you in a few decades...
Some are impressive:
- www.reddit.com/r/dalle2/comments/uzosy1/the_rest_of_mona_lisa
- www.reddit.com/r/dalle2/comments/vstuns/super_mario_getting_his_citizenship_at_ellis
And others are hilarious: - www.reddit.com/r/dalle2/comments/v0pjfr/a_photograph_of_a_street_sign_that_warns_drivers
- www.reddit.com/r/dalle2/comments/wbbkbb/healthy_food_at_mcdonalds
- www.reddit.com/r/dalle2/comments/wlfpax/the_elements_of_fire_water_earth_and_air_digitalhttp://www.reddit.com/r/dalle2/comments/uzosy1/the_rest_of_m...
http://www.reddit.com/r/dalle2/comments/vstuns/super_mario_g...
http://www.reddit.com/r/dalle2/comments/v0pjfr/a_photograph_...
http://www.reddit.com/r/dalle2/comments/wbbkbb/healthy_food_...
http://www.reddit.com/r/dalle2/comments/wlfpax/the_elements_...
http://old.reddit.com/r/dalle2/comments/uzosy1/the_rest_of_m...
http://old.reddit.com/r/dalle2/comments/vstuns/super_mario_g...
http://old.reddit.com/r/dalle2/comments/v0pjfr/a_photograph_...
http://old.reddit.com/r/dalle2/comments/wbbkbb/healthy_food_...
http://old.reddit.com/r/dalle2/comments/wlfpax/the_elements_...
https://www.reddit.com/r/dalle2/comments/v1sc2z/kermit_the_f...
I found this to be the most important point from this piece. Often people don't really know what they really want when it comes to creative work, let alone to some omniscient algorithm. In spite of that, it's a delight to see something you love from an unspecific prompt that you won't find with anything you receive from a human.
Dall.E 2 never ceases to amaze me.
For anyone interested in learning about what Dall.E 2 can do, the author also links to the Dall.E 2 prompt book (discussed in this post https://news.ycombinator.com/item?id=32322329).
That might be true, but after experimenting with DALL·E 2 last week (and spending more than $15), I have a different theory.
My tests focused on how well it could create art works around three common themes: still life, landscape, and portrait. For the first two categories, almost all the results were works that would not have looked out of place in a museum or art gallery. In contrast, with the prompt of “A painting of a young woman sitting in a chair” and variations, while DALL·E 2 produced convincing clothing, furniture, background, etc., the faces were mostly horrible. I started adding “from the rear” and “turned to the side” to the prompt just to get the face out of the picture.
I came to suspect that DALL·E 2 is bad at faces not because the developers made it that way but because human beings are uniquely hardwired to recognize faces. Most people are able to recognize and remember hundreds of faces, and we are very sensitive to minor changes in their configurations (i.e., facial expressions). When we look at a painting of a person sitting in a chair, we don’t care if aspects of the chair, the person’s clothing, etc. are not precisely accurate; a slight distortion of the face, however, can ruin the entire work. DALL·E 2 does not seem to have been trained to have the same sensitivity to faces that humans have.
If anyone is interested, the works that DALL·E 2 created for me are at [1]; video slideshows with musical accompaniment are at [2].
[1] http://www.gally.net/temp/dalleimages/index.html
[2] https://www.youtube.com/playlist?list=PLj4urky_8icRPzgFS_b98...
Dalle2 can clearly generate super-realistic faces without any problem, if you look at most of the posts at r/dalle2
The issue with small faces might be architectural if there is context-aware upscaling going on in the network, where a face needs to start larger than some smallest scale or it won't survive that process. That in turn might be an issue of too little training. A small face in a photo in the training data won't generate as much error gradient if it goes wrong as a larger face, but as you suggest we as viewers are much more prone to scrutinize faces even though they are small.
The reason it can't do faces well are very likely due to the filters being applied to try and stop people making pictures of real people. This is probably also the explanation for the random misses where it paints pictures of something that's not a llama. OpenAI is rewriting queries to make them more "diverse" i.e. acceptable to leftist ideology, and their rewriting logic seems to be completely broken. There have been many reports of people requesting something without even any humans in it at all, and discovering black/asian/arab people cropping up in it. At least earlier versions of the filter involved simply stuffing words onto the end as proven by people requesting "Person holding a sign that says " and getting back signs saying "black female" etc.
Man asks for a cowboy + a cat and gets a portrait of an Asian girl. Gwern comments with an explanation:
https://www.reddit.com/r/dalle2/comments/w7qvgl/comment/ihm6...
"tldr: it's the diversity stuff. Switch "cowboy" to "cowgirl", which would disable the diversity stuff because it's now explicitly asking for a 'girl', and OP's prompt works perfectly."
Big discussion thread where people discuss the problem and (of course) the censorship that tries to hide what's happening:
https://www.reddit.com/r/dalle2/comments/w944fa/there_is_evi...
"I once tried some food photography and received a cheese with a guys face for no reason."
"This has been mentioned on this sub multiple times, but those threads have consistently been removed by the mods - as will this one."
"There was a thread about that prompt and, yes, the person did get diverse [sumo wrestlers]"
"Been doing women images and seeing the article decided to try narrowing the results to "caucasian woman". Still gave me diversity. Whether you want it, or not, you're getting diversity"
Even when I re-used the exact prompts from the DALL-E Prompt Book, I didn't get anything near the level of quality and fidelity to the prompt that their examples did.
I know it's not a scam, because it's clearly doing amazing stuff under the hood, but I went away thinking that it wasn't as miraculous as it was claimed to be.
There is however one thing to be aware of, the titles posted on /r/dalle2/ and other places are often not the prompts that DALL-E2 got. Instead they are a fun description of the image done by a human after the fact. Random example:
"Chased by an amongus segway"
* https://www.reddit.com/r/dalle2/comments/wkv7za/chased_by_an...
But the actual prompt was:
"Award winning photo of a mole driving a red off road car through a field"
* https://labs.openai.com/s/xnaoxiWeSjiQX1QyVUCHGkl1
Which is quite a bit less impressive, as the actual prompt doesn't really match the image very well. And if you put "Chased by an amongus segway" into DALL-E2, you won't get an image of that quality either.
Here's a result for the prompt of "Woman with green skin, leaves instead of hair, wearing a simple dress, far shot, digital art, hyper-realistic, 8k, ultrahd," for example (all four images)
You will note that none of them are even basically fulfilling the prompt, as well as all four being, in my estimation, ugly and uninteresting. That's not unusual for prompts that involve some element of the fantastic -- though there are corners of less-realistic digital art that it does do well.
You will still note that none of them are far shots, that no depicted character actually has fully green skin, and one of the four has nothing even remotely like leaves for hair. I mean, is it better? Sure. They're less ugly, though none of them are what I'd call great results. But they also aren't really doing a basically competent job of fulfilling the prompt, much less producing a particularly striking or interesting images.
And my point is, outside of a few areas, this is what you get from Dall-E. Lots of misses, and if you're willing to put time into it and work on your results, a few hits. Don't get me wrong, I've gotten stuff from Dall-E that I think is great (I really like this "watercolor painting" for example: https://labs.openai.com/s/AQ7Wy5VHBWcLL5bJ5LbU5SuW) but I think it misrepresents Dall-E to suggest that most of the time it produces basically good images.
I'd say more like, "If you put time and attention into learning its quirks, in its best areas, it'll produce like one in ten images that are basically good."
And, I mean, on some level that's incredible. You can produce 10 images in about three minutes in Dall-E and get some great stuff. But I think people mostly see the top 10% of what Dall-E produces.
It's not fully removing humans from the equation, but you can take something that used to take days and make it a 20 minute operation.
I have my own api key as well but not with DALL-E 2 access just yet but seems similar in terms of prompting text in stages to get what you want. It feels kind of like negotiating with it in some way.
A lot of dreams scenery seems to throw logic and reasoning out of the window. Even small sensory inputs can make a huge difference to a dream sequence. And in many case they don't make sense even in the context of the dream.
I haven't personally experienced any hallucinations myself, but some DALL-E images seem awfully familiar to what some people describe.
I know that comparisons between brains and machine learning (including neural networks) are superficial at best, but I still wonder if DALL-E is mimicking, in its own way, a portion of our larger brain processing 'pipeline'.
Most people in our dreams don't even have faces that we would recognize. When they do have faces, sometimes it is not even the right face.
I love that we're at the level where the physical "realism" of correctly representing quadrupedals playing basketball is a thing now. I suppose the next level AI will be expected to model a full 3d environment with physical assumptions based on the prompt and then run the simulation
There's a lot of "80% there but not quite" in the current version, which makes it more of a novelty than a useful content generator.
The problem with moving to 3D is there are no almost no 3D data sources that combine textures, poses (where relevant), lighting, 3D geometry and (ideally) physics.
They can be inferred to some extent from 2D sources. But not reliably.
Humans operate effortlessly in 3D and creative humans have no issues with using 3D perceptions creatively.
But as for as most content is concerned it's a 2D world. Which is why AI art bots know the texture of everything and the geometry of nothing.
AI generation is going to be stuck at nearly-but-not-quite until that changes.
Film still of a llama in a jersey dunking a basketball like Michael Jordan, low angle, show from below, tilted frame, 35°, Dutch angle, extreme long shot, high detail, indoors, dramatic backlighting.
https://cdn.discordapp.com/attachments/999377404113981462/10...
Llama in a jersey dunking a basketball like Michael Jordan, screenshots from the Miyazaki anime movie
https://cdn.discordapp.com/attachments/999377404113981462/10...
The DALL-E 2 prompt book. If anything, pretty neat look at how the various prompts come out and some of the art created by it.
"How much are we paying him?"
"About $225k plus bonus and equity"
"And how much was the graphic designer paid?
"$55k"
"..."
Things that at one time took days can now be done in minutes with some skillful use of Dall-E + Photoshop. IMO, any image editing software that incorporates a similar technology will take over the market and it'll be one of the most important features in any graphic designer's toolkit.
A talented graphic designer who can also use a dall-e like tool is worth at least 5x the pay of one who can't (although I don't think we're going to get a "prompt engineer" title, it's really not that difficult a skill to pick up for people who already do image editing).
And we're still in the early days.
Unwillingly considering whether the easy bucks are worth the greasy feeling.
"I started out as a patty inversion engineer at McDonalds."
I mostly use it and Midjourney for material for my DnD campaign, but I'm going to need to do a little more work to make the whole thing coherent. Only tried it once and it was okay.
The interesting part is that it can do things like "female ice giant" reasonably whereas google will just give you sexy bikini ice giant for stuff like that which is not the vibe of my campaign!
Rather than trying to perfectly describe my image, I like to use references where the source material has what you want. With minimal direction these prompts get impressively close:
"larry bird as a llama, dramatic basketball dunk in a bright arena, low angle action shot, from the movie Madagascar (2005)" https://labs.openai.com/s/wxbIbXa0HRwwGUqQaKSLtzmR
"Michael Jordan as a llama dunking a basketball, Space Jam (1996)" https://labs.openai.com/s/mX4T5Iak8CMO1rPAmjRb7oyH
At this point I'd experiment with more stylized/recognizable references or add a couple "effects" to polish up the results.
I thought this was strange. Why hide an AI generated face?
Also, I came across this article which suggests that at some point users were not allowed to share images generating human faces, artificial or not: https://mixed-news.com/en/openais-dall-e-2-may-now-generate-...
Myself cant seem to get it to work. I think it's not very good at real things. Tried fitness related images, all is weird. Probably with fantasy kinda stuff its better since it has to be less accurate.
I think we're at the beginning of exploring what these image models can do and what the best ways to work with them are.
This is kind of funny. DALL·E is one of the most impressive pieces of software, but such a basic feature like history is curiously underpowered.
That’s not as easy as it sounds. Specially in the surreal cases that DALL-E is usually requested.
Sometimes you don’t know what you want until you see it. Other times you do, but are not able to express in ways that the computer can understand.
I see being able to communicate efficiently with the machine as a future in demand skill
Some guy spent hours feeding the AI pictures he liked to get an end result he was happy with.
It needs a better text to image model, I think. Maybe you can fork it and improve?
https://docs.google.com/presentation/d/1y8EE_p8bw9dIEDguT1bT...
The fact that it's a derivative of an existing work is noteworthy, but I gave it absolutely no guidance on the topic. If i suggest something it will give it a go with similar fervor. eg https://imgur.com/a/N1qWaSV
I ran out of credits way too fast, so I like to see other people playing with it and their iterative process.
Sounds awfully like programming...
I've been playing with Stable Diffusion a lot, and in my experience its results are much weaker then what's shown in this post. The artistic pictures that it generates are beautiful, often more beautiful then Dalle-2 ones. But it has a real problem understanding the basic concepts of anything that is not the simplest task like "draw a character in this or that style". And explaining the situations in detail doesn't help - the AI just stumbles upon basic requests.
Seems like Stable Diffusion has a much more shallow understanding of what it draws and can only produce good result for things very similar to the images it learned from. For example, it could generate really good dutch still life paintings for me - with fruits, bottles and all the regular expected objects for this genre of painting. But when I've asked it to add some unusual objects to the painting (like a Nintendo switch, or a laptop) - it couldn't grasp this concept and just added more warbled fruit. Even though the system definitely knows how a Switch looks like.
The results in the post are much more impressive. I doubt that Dalle-2 saw a lot of similar images in training, but in all of the styles and examples it definitely understood how a llama would interact with a basketball, what are their relative sizes and stuff like that. On surface results from different engines might look similar, but to me this is an enormous difference in quality and sophistication.
I'm more curious of how this will effect stock photography. Soon anyone can generate the exact image they're looking for, no matter how obscure.
I wonder what Gary Marcus or Filip Pieknewski think about it. Surely they must be eating crow.
What wrote the prompt?
And this AI doesn't. Your anecdote is totally unrelated to the idea of AGI in the gp post. The fact that it made you laugh is a happenstance. It was not "trying" to make you laugh.
I also saw this one recently from Midjourney. Would not call the humor random.
https://www.reddit.com/r/midjourney/comments/w73rhv/prompt_t...
It’s still up on the dalle2 subreddit.
I suspect AGI, depending on how its defined, will be with us in some form in the next few decades at most. Just a hunch. This is nothing to do with that mission though imho. Maybe you can read into it something like, "we are solving lots of discrete problems like this, maybe we can somehow glue them together into a higher level program"? That might give you something AI-esque? My guess is that 'true' AGI will have an elegant solution rather than a big bag of stuff glued together.
An AGI wouldn't need us to this extent, or at all. An AGI would also be able to come up with new ways to represent ideas, even ways that are foreign to us.
We are not. But maybe we are closer to replicating some of our internal brain workings.