Or if you want to see some in 3D, OpenAI featured a pelican riding a tricycle, bicycle, pony and another pelican in their livestream this morning: https://www.youtube.com/live/Wq45rvPGNHs?t=1070s
Or if you want to see some in 3D, OpenAI featured a pelican riding a tricycle, bicycle, pony and another pelican in their livestream this morning: https://www.youtube.com/live/Wq45rvPGNHs?t=1070s
On the other hand: the test is clearly not saturated, given that you can see a clear difference in output at the various reasoning levels / model versions.
https://themagnet.substack.com/p/why-is-it-so-hard-to-draw-a...
But either way, with no real way to visualize the result of the text it starts with - it will always be stabbing in the dark. It can't understand conceptually what any of it should look like and then refine the SVG to improve it gradually. It just throws darts at a wall and hopes it comes out alright.
Have them use tikz instead of svg, or have it write code that moves the cursor and draws the thing in paint.
Compositionality and visualization are generally much, much better at each new generation / release cycle.
It's fascinating how well models have internalized visualizing things without actually having joint embeddings / broad multimodality.
https://chatgpt.com/share/6a5009de-fff8-83ea-98ff-0da17d1d04...
I once used something like karpathy's auto-scientist to mutate the prompts and rank them with a vison model. Some of the winners where pretty neat. I think they have a lot more style than the gpt-5.6 ones. https://xcancel.com/xundecidability/status/20449185674144196...
A skilled human artist wouldn't have both legs in front of the bike, or a single straight line representing both leg's crank arms.
Dead internet theory? Semi-random parroting by real people? Or something else.
I assume multimodal models can do it already do it today if constantly asked "make it better"
Also would be good to have a tool where users can select models and instantly see each model's generated pelicans. That will make it easy to compare the output of different models.
https://www.google.com/search?client=firefox-b-d&q=pelican+1...
Nice to see you did the quality level comparisons and did three passes.
I've been using that technique myself on my image gen reviews[1] and it also works well in presentations and for personal study.
[1] https://generative-ai.review/2025/12/beast-mode-activated-op...