Exactly this. At a high level, DALL-E is mapping text to a (continuous) matrix and then mapping that matrix to an image (another a matrix). All text inputs will map to _something_. DALL-E doesn't care if that mapping makes sense, it has been trained to produce high-quality outputs, not to ensure the validity of mappings.
None of this makes DALL-E any less impressive to me. High quality image generation is a truly amazing result. Results from foundational models (GPT-3, PaLM, DALL-E, etc) are so impressive that they're forcing us to reconsider the nature of intelligence and raise the bar. That's a sign of a job well done to me.