I dunno man, I punched that exact prompt ("man throwing his smartphone into a river") in DALL-E 2 just now, and in 2/4 samples, the smartphone is clearly separate from the hand: labs.openai.com/s/uIldzs2efWWnm3i9XjsHI7or labs.openai.com/s/jSk4qhAxSiL7QJo7zeGp6m9f
> The semantics don't seem to be fully formed.
Yes, not so much 'formed' as 'formed and then scrambled'. This is due to unCLIP, as clearly documented in the DALL-E 2 paper, and even clearer when you contrast to the GLIDE paper (which DALL-E 2 is based on) or Imagen or Parti. Injecting the contrastive embedding to override a regular embedding tradesoff visual creativity/diversity for the semantics, so if you insist on exact semantics, DALL-E 2 samples are only a lower bound on what the model can do. It does a reasonable job, better than many systems up until like last year, but not as good as it could if you weren't forced to use unCLIP. You're only seeing what it can do after being scrambled through unCLIP. (This is why Imagen or Parti can accurately pull off what feels like absurdly complex descriptions - seriously, look at the examples in their papers! - but people also tend to describe them as 'bland'.)