- Describe the task creation process, why they are appropriate for measuring X, etc. Preferably drawing from similar studies in people instead of ad hoc. Lots of work has already gone into measuring these things. They may have to be modified for DALL-E but it would be a better starting point.
- Show several variations of the same task and outputs. For example #11, maybe the issue is that DALL-E has a poor understanding of milk sizes as expressed but not size relationships in general. Example #13 with pizza sizes is ambiguous by the authors own interpretation yet they deem it a failure. It would be trivial to construct dozens of similar examples to give a more holistic understanding.
- Narrow the scope of the prompts if you want to see how the model understands a particular relationship. Many of the tasks include multiple ancillary statements.
- Discuss prompt engineering in more depth. We already know that these models are sensitive to the formulation of the prompt. What did the authors try / not try?
- Replace the authors' individual opinions of the outputs with crowdsourced opinions from mturk or even Twitter. As I noted above, example #10 is not as clear cut as the authors suggest.
- Measure how well people do at the same task as a baseline: draw something that aligns with a given text prompt and compare to DALL-E, maybe with a crowdsourced opinion of which is a more accurate interpretation. Even if they're just stick figures this would be interesting to see.
As-is, this article doesn't really add anything substatial about the model's capabilities to the conversation.
The limitation is caused by the CLIP model they used to encode text and images. It's a separate model only generating an embedding, it's not using attention and pairwise interactions on the whole sequence. This causes Dalle2 to be bad at handling multiple objects with multiple attributes. There is no reason the complete prompt could not be related to the generated image instead of an embed, thus correctly stacking the coloured cubes and assigning the right age to each person mentioned in the prompt.
please actually read my work and please don’t make stuff up.