> To the extent that the goal is to develop artificial intelligence that can be trusted in safety-critical applications (Marcus & Davis, 2019), a much higher standard must be applied.
I have seen no mention of that being the intended use of DALL-E (2). In fact, the most common use case I've seen described is in replacing Fiverr-type tasks: quick graphic design.
That said, the results are actually encouraging given that it _wasn't_ designed to succeed here:
> Nevertheless, for 5 out of the 14 prompts, at least one of the ten images fully satisfied our requests.
Some of the authors' interpretations could be argued against, as well. For instance, in example 10, "An old man is talking to his parents":
> In none of these images did DALL-E successfully infer that image should show an old man with two even older people
Several of the images appear to show exactly that? How is the author judging "even older"?