I've not tested this, but I suspect you can get DALLE to create pictures that humans are more likely to describe as 'nonsense' by adding 'nonsense' or 'nonsensical' to the prompt. That'd indicate that it absolutely
does have an idea about 'nonsense' and can recognize, & reproduce within its constrained output, 'nonsense' that's largely compatible with human ideas of what 'nonsense' is.
Asking it to produce noise, or raise an objection that a prompt isn't sufficiently meaningful to render, is a silly standard because it's been designed, and trained, to always give some result. Humans who can object have been trained differently.
Also, the GPT models – another similar train-by-example deep-neural architecture – can give far better answers, or give sensible evaluations of the quality of its answer, when properly prompted to do so. If you wanted a model that'd flag nonsense, just give it enough examples, and enough range-of-output where the answer your demanding is even possible, and it'll do it. Maybe better than people.
The circumstances & limits of the single-medium (text, or captioned image) training goals, and allowable outputs, absolutely establish that these are different from a full-fledged human. A human has decades of reinforcement-training via multiple senses, and more output options, among other things.
But to observe that difference and conclude these models don't "understand" the concepts they are so deftly remixing, or are "just a very complex and impressive mirror", does not follow from the mere difference.
In their single-modalities, constrained as they may be, they can train the equivalent of a million lifetimes of reading, or image-rendering. Objectively, they're arguable now better at composing college-level essays, or rendering many kinds of art, than most random humans picked off the street would be. Maybe even better than 90% of all humans on earth at these narrow tasks. And, their rate of improvement seems only a matter of how much model-size & training-data they're given.
Further: the narrowness of the tasks is by designers' choice, NOT inherent to the architectures. You could – and active projects are – training similar multi-modality networks. A mixed GPT/DALLE that renders essays with embedded supporting pictures/graphs isn't implausible.