> Whether results of this kind should be considered as successes for the program – what is the proper measure to use in evaluating success – depends on the intended use of the program. If the goal is to generate candidate images that a graphic artist will choose from, or choose from and edit, then the system can reasonably be measured in terms of the quality of the best result out of ten or out of one hundred.
They basically admit your Fiverr use case is valid. But say, that it should not be used "in safety-critical applications" which is neither a grand claim nor controversial. It is probably the most blasé claim because, as you point out, no one is expecting this to be used in safety-critical applications. From an economics point of view, the Fiverr use case seems pretty strained to me. If you've ever watched street art, some dazzling things can be done in under 10 minutes. Unless the DALL-E gets it correct on the first shot +99% of the time, someone sifting through images is probably just as costly as paying for Fiverr. What this paper elucidates to me is that even historical figures are off-limits which, in my expectation, is a non-trivial use case.