None of this makes DALL-E any less impressive to me. High quality image generation is a truly amazing result. Results from foundational models (GPT-3, PaLM, DALL-E, etc) are so impressive that they're forcing us to reconsider the nature of intelligence and raise the bar. That's a sign of a job well done to me.
As much as people would like there to be, there really does not seem to be anything here. The original author doesn't think so, either (would need to refind the tweet).
Just nitpicking here a bit. It has been trained to ensure the validity of mappings, but only for mappings of valid prompts, where "valid" is vaguely "things that appear in the training set". On the other hand, it wasn't trained to ensure the validity of mappings of invalid prompts to images. It's an open question what that would even mean - someone in another thread here suggested it should output "not a valid prompt" in this case.
This is by far the worst.