Given the model has only 1 input and 1 output and training is essentially surviving that order, it's not dissimilar.
This is probably because your string is a low-entropy keyboard-mash.
Of course. And the equivalent of that explanation is baked into DALL-E, in the form of its programming to always generate an image.
> but DALLE can not act like humans
No, not generally, but I don't think anyone has claimed that.
So is a single-purpose AI equivalent to the entirety of the Human Experience? Of course not. But can it be similar in functionality to a small sliver of it?
One great example of this phenomenon is "Häagen-Dazs".[0]
Admittedly that's a brand name, rather than a specific dish, but I assume that Dalle-2 would generate an image of ice cream if given a prompt with that term in it (unless there is a restriction on trademarks?).
[0] https://funfactz.com/food-and-drink-facts/haagen-dazs-name/
Asking it to produce noise, or raise an objection that a prompt isn't sufficiently meaningful to render, is a silly standard because it's been designed, and trained, to always give some result. Humans who can object have been trained differently.
Also, the GPT models – another similar train-by-example deep-neural architecture – can give far better answers, or give sensible evaluations of the quality of its answer, when properly prompted to do so. If you wanted a model that'd flag nonsense, just give it enough examples, and enough range-of-output where the answer your demanding is even possible, and it'll do it. Maybe better than people.
The circumstances & limits of the single-medium (text, or captioned image) training goals, and allowable outputs, absolutely establish that these are different from a full-fledged human. A human has decades of reinforcement-training via multiple senses, and more output options, among other things.
But to observe that difference and conclude these models don't "understand" the concepts they are so deftly remixing, or are "just a very complex and impressive mirror", does not follow from the mere difference.
In their single-modalities, constrained as they may be, they can train the equivalent of a million lifetimes of reading, or image-rendering. Objectively, they're arguable now better at composing college-level essays, or rendering many kinds of art, than most random humans picked off the street would be. Maybe even better than 90% of all humans on earth at these narrow tasks. And, their rate of improvement seems only a matter of how much model-size & training-data they're given.
Further: the narrowness of the tasks is by designers' choice, NOT inherent to the architectures. You could – and active projects are – training similar multi-modality networks. A mixed GPT/DALLE that renders essays with embedded supporting pictures/graphs isn't implausible.