If DALL-E just spits out 1000 images and then a human goes through them and picks the best 2-3, and those are good - it's impressive, but the human was still a crucial part of the process. On the other hand, if DALL-E were to generate 10 billion images, and choose the best 2-3 itself and give those as output, and if at least one of those 2-3 would be consistently great, then DALL-E could be indeed considered to be creating (good) art.
It's worth noting that the OpenAI samples for DALL-E 1 used CLIP to rank generated samples, and got a big boost from that. For many model architectures, you can run them in reverse to do 'image -> caption', and 'score the caption' quality: if 'the caption is bad', that indicates your image was screwed up and low-quality (introduced by Cogview). DALL-E 2 doesn't use either approach, or finetuning on user choices like InstructGPT, and I dunno if OA is going to implement any of these, but there is a wide universe of techniques applicable here to improve quality and we should keep that in mind (https://www.gwern.net/Forking-Paths) if we are going to make any assertions more sweeping than "this specific model, at this very instant, with this particular interface, is only at this level of quality".
I disagree. The human artist's tastes at least partially originate in other people, both individuals and general societies/cultures, and oftentimes the artist directly incorporates feedback into future work. Are you aware that students in art school, music conservatories, etc constantly get feedback from instructors and peers? I reject your premise entirely unless you can give me an example of a human that created art without ever having been influenced as a human being by any other human being. Otherwise I believe it's just what I said before: concluding first that AI can't create art and finding reasons second.
Until then, my point remains: DALL-E is currently like an (extraordinarily good) hat that you can put words in and extract phrases out of. A human chooses what words to put in and which of the phrases they take out are better. Unlike pulling words out of a hat, the network has some criteria by which it produces phrases, but that's not enough to call it an artist.
This is not meant to minimize how good the achievement of this network is. The level of fidelity and even understanding of the prompts is extraordinary. But its purpose is not to be creative, it is to find a point on a hyperplane that matches the input it received. It is currently at the level of a tool - though there are potential advancements that could yet turn it into an artist in its own right.
When DALL-E x.0 does that, and when it also generates similar quality from much higher-level prompts ("paint a sad picture", or "social commentary on BLM" or something like this, instead of a description of what the picture should show and in what style), then I for one will be in complete agreement that it's indeed an artist in itself.
Personally, I don't expect this to happen in the next few decades, as I don't think the current approaches are very promising for the type of intelligence that you would need to actually do this type of reasoning, but that remains to be seen, and I am fully confident that it will happen some day.
Fairly certain you can get results using this prompt with no issues.
Dalle 1.0 -> 2.0 was a shocking improvement, I expect the jump to 3.0 to be equally jarring if not more so
Van Gogh invented Starry Night without any prompting despite it not being a real scene (much less anything he had ever seen and such abstraction was very rare in the 1880s). Picasso made Les Demoiselles d’Avignon in 1907; it was so radical even his fellow artists were unable to comprehend it.
It doesn't change the fact that DALL-E is pretty amazing tech, but it's still as far behind human ability as any AI is today. It is way way better than what came before, but that's true of most technologies.
It's just like Dada poets did 100 years ago: you use a mechanical process to generate quasi-random output, and then you choose some of this output to present to other humans. The way you provide input to the mechanical process (e.g. what words you choose to put in the hat) and the curation are the real creative part of the art, not the mechanical process generating the text itself.
For this one in particular, here are a few more results for Battlestar and The Office:
https://twitter.com/Miles_Brundage/status/153247388947686195...