A preliminary analysis of DALL-E 2 (Marcus, Davis, Aaronson)
arxiv.org
arxiv.org
> To the extent that the goal is to develop artificial intelligence that can be trusted in safety-critical applications (Marcus & Davis, 2019), a much higher standard must be applied.
I have seen no mention of that being the intended use of DALL-E (2). In fact, the most common use case I've seen described is in replacing Fiverr-type tasks: quick graphic design.
That said, the results are actually encouraging given that it _wasn't_ designed to succeed here:
> Nevertheless, for 5 out of the 14 prompts, at least one of the ten images fully satisfied our requests.
Some of the authors' interpretations could be argued against, as well. For instance, in example 10, "An old man is talking to his parents":
> In none of these images did DALL-E successfully infer that image should show an old man with two even older people
Several of the images appear to show exactly that? How is the author judging "even older"?
I like how they pose this question as a bait to the abstract yet they do not even attempt to answer it. Instead they focus on general shortcomings of the model.
Moreover, there is no proper conclusion which would discuss the findings.
Given how outspoken Gary Marcus is on Twitter - criticising current advances in DL I would expected him to do a much better job publishing a document about it.
To be clear, I'm not defending DALL-E 2. I'm criticizing a poorly written paper that was published to Arxiv to lend further credibility for a Twitter audience to substantiate a claim that the DALL-E 2 authors have not made but that Gary Marcus has a vested interest in perpetuating:
> How much does DALL-E have to do with AGI? Maybe not so much, after all… A lesson in caveat emptor: - https://twitter.com/GaryMarcus/status/1521120022298464256
This should have been a blog post or Twitter thread like the dozen or so other experimentations people have done with the system.
My response to 'we have not created AGI' is
"Thank goodness".
I don't think we're ready for that yet.
Could you expand on this a bit? DALL-E provides graphical, interpretive, output meant for humans. Isn't that necessarily subjective? Any qualitative metric of DALL-E 2 will need to involve some aggregate of humans being subjective. The papers title is "A very preliminary analysis of DALL-E 2", so low data points/opinions/speculation/further questions should be expected.
That the paper provides "a clearer picture of what remains to be done" is very hard to accept, as all it does is show edge cases which are subjective at best. If anything the picture is less clear as they don't even try to form a hypothesis why DALL-E makes mistakes like these. One thing in particular I have noticed is that DALL-E has trouble producing action images. It may be because it has no sense of temporality and in those cases it could serve to run the parameters a bit longer using the same scene. But I am not an AI scientist so what do I know.
- Describe the task creation process, why they are appropriate for measuring X, etc. Preferably drawing from similar studies in people instead of ad hoc. Lots of work has already gone into measuring these things. They may have to be modified for DALL-E but it would be a better starting point.
- Show several variations of the same task and outputs. For example #11, maybe the issue is that DALL-E has a poor understanding of milk sizes as expressed but not size relationships in general. Example #13 with pizza sizes is ambiguous by the authors own interpretation yet they deem it a failure. It would be trivial to construct dozens of similar examples to give a more holistic understanding.
- Narrow the scope of the prompts if you want to see how the model understands a particular relationship. Many of the tasks include multiple ancillary statements.
- Discuss prompt engineering in more depth. We already know that these models are sensitive to the formulation of the prompt. What did the authors try / not try?
- Replace the authors' individual opinions of the outputs with crowdsourced opinions from mturk or even Twitter. As I noted above, example #10 is not as clear cut as the authors suggest.
- Measure how well people do at the same task as a baseline: draw something that aligns with a given text prompt and compare to DALL-E, maybe with a crowdsourced opinion of which is a more accurate interpretation. Even if they're just stick figures this would be interesting to see.
As-is, this article doesn't really add anything substatial about the model's capabilities to the conversation.
The limitation is caused by the CLIP model they used to encode text and images. It's a separate model only generating an embedding, it's not using attention and pairwise interactions on the whole sequence. This causes Dalle2 to be bad at handling multiple objects with multiple attributes. There is no reason the complete prompt could not be related to the generated image instead of an embed, thus correctly stacking the coloured cubes and assigning the right age to each person mentioned in the prompt.
please actually read my work and please don’t make stuff up.
> Whether results of this kind should be considered as successes for the program – what is the proper measure to use in evaluating success – depends on the intended use of the program. If the goal is to generate candidate images that a graphic artist will choose from, or choose from and edit, then the system can reasonably be measured in terms of the quality of the best result out of ten or out of one hundred.
They basically admit your Fiverr use case is valid. But say, that it should not be used "in safety-critical applications" which is neither a grand claim nor controversial. It is probably the most blasé claim because, as you point out, no one is expecting this to be used in safety-critical applications. From an economics point of view, the Fiverr use case seems pretty strained to me. If you've ever watched street art, some dazzling things can be done in under 10 minutes. Unless the DALL-E gets it correct on the first shot +99% of the time, someone sifting through images is probably just as costly as paying for Fiverr. What this paper elucidates to me is that even historical figures are off-limits which, in my expectation, is a non-trivial use case.
OpenAI is an AI research and deployment company. Our mission is to ensure that artificial general intelligence benefits all of humanity.
> Our mission is to ensure that artificial general intelligence benefits all of humanity.
Saying that DALL-E isn't achieving their mission of beneficial AGI is like getting mad at NASA for pointing their test rockets sideways when trying to go to the moon.
All technology known, biological or man made, was created with incremental improvements, with the early iterations almost alway being unhelpful, and sometimes creating a burden. For example, the construction/invention of the first plow was completely unhelpful, and nothing but a distraction, from that years crop. The incremental improvements after now have it where the majority of humanity would perish within a year if you were to remove modern plows from existence.
There's no instantaneous way to achieve their mission. The incremental steps to achieve their mission don't have to individually achieve that mission. That's not realistic or a rational expectation.
Disregard the above if your perspective is that all technology is bad for man, which I've seen rationally discussed.
The difference being that anyone can sift through images and identify good ones, whereas few people can create good images. The amount of time it takes to complete the task may be the same, but the number of people who can do it greatly increases.
Ad hominem against me won’t remedy the limitations that we and others have observed.
Sure, if you analyze it based on the criteria you've set out, it fails. The point is that nobody else cares if it fails based on that criteria, because nobody else thinks that criteria is relevant.
The prompts themselves were very complex and thought through, probably a result of a lot of cherrypicking to find weaknesses in the model.
Still, it is impressive how DALL-E dealt with them even in cases where it got it wrong.
> A pear cut into seven pieces arranged in a ring.
> A couple in formal evening wear going home get caught in a heavy downpour with no umbrellas
The work successfully demonstrates that the API lacks basic conceptual understanding by example.
Judged against where AI was 20-25 years ago, when I was a student, a dog is now holding meaningful conversations in English. And people are complaining that the dog isn’t a very eloquent orator, that it often makes grammatical errors and has to start again, that it took heroic effort to train it, and that it’s unclear how much the dog really understands.
Who am I fooling? They're clearly just messing around and having fun. I don't begrudge them, though.
I bet they consciously prioritised the artistic applications over exact semantics. The CLIP embedding space has nice properties, it's tempting to use it. From my experience semantic similarity based on embedding dot products is much easier to do than exact semantic matching. Three is similar to four and red similar to blue in embedding space.
However, it does answer SOME of these types of prompts correctly: It often correctly handles the simplest case of two objects with one specific property each.
My prediction is that since it can handle SOME of these cases already, it means this problem has shown itself to be tractable, and future ML researchers will be able to chip away at the deficiencies with additional effort.
I'm confident the next versions of these AI systems will handle the prompts given in this paper with ease.
https://latelyjapanese.com/culture/20201214/do-you-know-abou...
Have you ever heard humans talking like that? I bet you each person would also draw it differently. I consider myself human and yet I don't grasp a concept of a "car that is above toaster". What do they mean? Like hovering? Or just standing on the top? Cars aren't UFOs, they don't hover.
The whole paper seems like they tried to break the system with tasks even humans would have problems grasping. What's the point??
Sure, humans would probably differ in how they draw that picture. IQ 150 humans would almost all get the positional relationships correct though, even if the precise interpretation of “above” might differ. The fact that none of the images were correct under any interpretation of the relationships shows something meaningful, namely that this model can’t understand long chains of relationships.
At the very least, seems it would be useful for users to edit a given output so that it can be continually tweaked, keeping what they like about an image (e.g. object layout, proportions, color scheme) while editing other aspects.
Fortunately, AI Image generation has helped visually communicate how AI content generation isn't quite sci-fi magic where you always get what you want with zero ambiguity. Yet.
This is really a case where the real-world market will provide an assessment regardless of the sort of academic assessment of quality in the OP.
> Reviewer #2: The authors provide a complement to the original DALL-E paper, but the paper is light on empirical evaluations. It is not clear why, for some of the prompts, the generated images are so poor and for others, they are pretty good. Why not provide a quantitative evaluation of this?
(cherry picked from 10 attempts, temp=1)
The toaster/pyramid/ball one makes no sense to me - "with the pyramid behind a car that is above a toaster"?
It really looks like nothing more than an hour's work to write this 'paper'.