Much Win! ;)
This relies on the false premise that, if they would include it in their training dataset, it would be perfect. All they need to do is be good enough and better than the other, not perfect.
Based on the one Simon commented though, I'd say we're in decent territory to try the latter part of his hypothesis.
In all seriousness, that's what makes it an interesting test: it's asking for something technically impossible, that requires artistic license to make coherent.
Making specific choices on where to bend reality (and where not to) is a big chunk of visual art.
I mean the prompt was succinct and clear, as always - and it still decided to hallucinate multiple features (animation + controls) beyond the prompt.
It'd also like to point out that to date no drawing was actually good from an actual quality perspective (as in comparative to what a decent designer would throw together)
Theyre always only "good" from the perspective of it being a one shot low effort prompt. Very little content for training purposes.
And so if you ask it to do something big it will do a very surface level implementation. But if you have it iterate many times, or give it small pieces each time, you’ll end up with something closer to what a human would do.
I imagine the pelican test but done in a harness that has the agents iterate 10+ times would be closer to what you’d expect, especially if a visual model was critiquing each time.
It would always look goofy - by design, but it usually looked good.