To be clear, I'm not defending DALL-E 2. I'm criticizing a poorly written paper that was published to Arxiv to lend further credibility for a Twitter audience to substantiate a claim that the DALL-E 2 authors have not made but that Gary Marcus has a vested interest in perpetuating:
> How much does DALL-E have to do with AGI? Maybe not so much, after all… A lesson in caveat emptor: - https://twitter.com/GaryMarcus/status/1521120022298464256
This should have been a blog post or Twitter thread like the dozen or so other experimentations people have done with the system.
My response to 'we have not created AGI' is
"Thank goodness".
I don't think we're ready for that yet.
Could you expand on this a bit? DALL-E provides graphical, interpretive, output meant for humans. Isn't that necessarily subjective? Any qualitative metric of DALL-E 2 will need to involve some aggregate of humans being subjective. The papers title is "A very preliminary analysis of DALL-E 2", so low data points/opinions/speculation/further questions should be expected.
That the paper provides "a clearer picture of what remains to be done" is very hard to accept, as all it does is show edge cases which are subjective at best. If anything the picture is less clear as they don't even try to form a hypothesis why DALL-E makes mistakes like these. One thing in particular I have noticed is that DALL-E has trouble producing action images. It may be because it has no sense of temporality and in those cases it could serve to run the parameters a bit longer using the same scene. But I am not an AI scientist so what do I know.
I like how they pose this question as a bait to the abstract yet they do not even attempt to answer it. Instead they focus on general shortcomings of the model.
Moreover, there is no proper conclusion which would discuss the findings.
Given how outspoken Gary Marcus is on Twitter - criticising current advances in DL I would expected him to do a much better job publishing a document about it.