I'm actually beginning to wonder if some people who ignore these things have a different, somewhat lesser ability to percieve image details than I do.
I mean I guess its fine to go on to another test despite never actually passing the pelican bike test, but there's a sense that we have to use another test because AI is now good at pelicans on bikes, which is just not true.
Please remember, we've started from there :
https://simonwillison.net/2024/Oct/25/pelicans-on-a-bicycle/
When it started, it was clear what LLM would stand out, its style, etc. Nowadays, the pelicans look similar, the difference is in details and sometimes hard to catch. Sure, the task is not completed perfectly, but that's not the point. It was supposed to be a benchmark to quickly benchmark a LLM against others.
Isn't it?
If the computer can't do it better than a human being, then what's the point?
Being wrong at scale is not better than being right.
But no ones hire random humans for things like this. You go and hire a vector artist and they will get your a very good pelican on a a bike. That's how you get things done when you can't do it.
Yeah, but then, you recruit the artist for $XXX - whereas you "recruit" your LLM for $0.XXX for the same task.
Of course the quality difference is huge. But sometimes you don't need that level of quality.
Also, finding a vector artist takes days of communication, payment settlement, revisions, etc.
Not always the most practical solution.
Because the benchmark wasn't testing "can an LLM draw a pelican like a human". The original article was testing the relative capabilities between LLMs. Now that LLMs can all draw pelicans all similarly, the test is less interesting as a comparative benchmark.
This is what the tech industry has become?
Less of a failure is still failure.
LLM has progressed a lot in the last two year, judging from the pelican drawings. I personally couldn't care less about it though. I do know that I've gone from using no AI at all for coding to probably 95%. I hardly code by hand anymore. That's much more impressive and significant. Failure you said?
Pick up a newspaper. Start with the Wall Street Journal. These are public companies. It's not a secret.
It can certainly do it better than I can. Sometimes you don't have a human handy with the required skills to do something.
It's physically impossible.
The problem is to draw it in the least disturbing way possible.
https://ridesabike.com/donald-duck-daisy-duck-huey-dewey-and...
When you aren't sure if an LLM can write an svg well, or that it will be able to form a pelican shape, or animate a bicycle, it's a good test. After that, it's all judgement: how detailed should the pelican be? pelicans are the wrong shape for a bicycle by default, so how much can I change its physiology to match using a bicycle before it isn't a pelican? Do I care about how well the client is able to render complex geometry?
It's not that there isn't room to do better, or that it doesn't tell you anything at all, but rather we've reached a point where what it tells us isn't very clear anymore.
If the expectation is that AI is going to replace "knowledge workers" then the limit would be a darn perfect drawing. We are nowhere close to that.
And Elon is already propagating the age of abundance where money won't exist anymore, right before calling the interviewing journalist dishonest and deservedly losing public trust. Smh my head.
What knowledge workers do you know that have excellent drawing skills? I worked in a design agency and for a couple of years, each week me and a few other people would attempt to sketch a member of our group: one person would be the model and sit still, and everyone else would draw her/him.
Let me tell you, if producing a convincing portrait was a prerequisite for being a knowledge worker, there would be 99% fewer knowledge workers.
It surprises me how many people in this community don't get this. Obviously, most prompts thrown into an AI chatbot/interface are about something no knowledge worker would ever have to deal with. That doesn't disqualify them as a tool for measuring progress of the models.
Oh come now. I am extremely confident that if I hired a professional artist to draw a picture of a pelican riding a bicycle, I would get something inarguably much better than what today's best coding LLMs can produce.
I think AI folks have done a terrible job of communicating this, but replacing a professional simply isn't the point. The point is to serve all the situations where people would've never considered hiring a professional, and where perfection or artistic merit isn't the point (say a personal throwaway recreation of an LOTR world).
And I think in that regard the benchmarks are pretty good.
I'm not saying it is, just that there's obviously still room for the models to improve on this task.
* some omitted the bottom of the diamond which connects from the pedals to the rear wheel
* some added an extra connection from the pedals to the front wheel, making it impossible to steer
* none could align the head tube with the fork
* none added a correct offset to the fork
* none could generate the chain properly in a way that attaches to the two sprockets correctly
I mean just look at these:
* Grok 4.5: https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/p...
* GPT 5.6 Terra: https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/p...
* Sonnet 5: https://s3.eu-west-1.amazonaws.com/images.dylancastillo.co/p...
The bike is generally okay, apart from medium which derped hard. Max has correct diamond, correct head tube, and so on. Only nitpicks are that the front fork offset isn't there and the chain doesn't touch the rear sprocket correctly.
It used to be a very difficult task for models, see [2,3,4]
it cuts across several tasks that AI used to be very bad at, but now has improved quite a bit. Namely, spatial reasoning (because it has to manually place the points of the svg such that they make sense and form what it says it forms. This used to not work very well, with random shapes floating around that it would mark things like "eyebrows" but were nowhere near the "eyes", etc.
It also tests the model's world knowledge (what do pelicans look like? sure they have wings, feet, beaks etc, but what shape are they? how to get proportions roughly right? this isn't a given from text data about the bird. This goes doubly for a bike, which is a quite complex shape that most humans fail to draw correctly[1] (many draw the frame or chain connecting in impossible ways that would not ever function mechanically)
Before it was pelican on a bicycle there were people having it do horses/unicorns making the rounds - gpt4.0 or whatever would often make hideous abominations of legs and mouths
[1] https://www.gianlucagimini.it/portfolio-item/velocipedia/
[2] https://static.simonwillison.net/static/2026/mistral-small-4...
[3] https://static.simonwillison.net/static/2025/codex-hacking-m...
[4] https://static.simonwillison.net/static/2025/gemini-2.5-flas...
So it's a pretty much pointless test now.
It's quite surprising actually.
Not only that this so-called "benchmark" isn't economically useful, but that it tests for the sake of testing; and for attention.
To end this obsession with generating SVG pelicans, Quiver AI's [0] model is actually designed to generate SVGs from prompts and has done so for years.
So there is no need to continue with this un-serious pseudoscientific "benchmark".
Does that even test for intelligence?
This is like testing if a horse can fly just because someone showed an image of a Pegasus, or testing if a fish can climb up a tree and believing they are not intelligent because each of them cannot fly or climb up trees.
Pulling things into the ridiculous isn't condusive to a good faith discussion and does not help proving pseudoscientificness which you seem to be after. (I don't belive that the pelican bicycle test is meant to be a serious scientific endeavour, btw.)
Then use a model that supports multi-modal output then, instead of using models that only support raw text output to construct an image.
> If it can't even do that well then I'm not sure why we need to talk about horses here.
Testing a model that only outputs text and forcing it to output an unsupported modality is obviously not going to work well which is what it going on here.
Would you expect Claude/GPT to replace your compiler because it can output text? That would be like expecting that horses can fly because someone saw an ancient winged-horse on a brick tablet.
> Pulling things into the ridiculous isn't condusive to a good faith discussion and does not help proving pseudoscientificness which you seem to be after.
That is because the premise of the test is ridiculous and pseudoscientific.
> (I don't belive that the pelican bicycle test is meant to be a serious scientific endeavour, btw.)
Why did you say it was a "test for intelligence", when we both know it is a joke that tests for nothing?
Somebody who has trained painting animals and vehicles for decades is surely skilled at that, but is not solving new problems, so is not necessarily applying intelligence when painting something of that sort.
You can take a reasonably skilled programmer, show them some C code and ask them to write equivalent ASM code, perhaps providing them with a language reference for either, and they will be able to do this translation. Slowly, but they will get there. They have not been trained to do this, but they will "figure it out". With enough time and motivation, even non-programmers would be able to do that. That is what intelligence is and of course Claude/GPT needs to be able to do that if it wants to claim that it intelligent. No flying horses.