> give me the svg of a pelican riding a bicycle
> I am sorry, I cannot provide SVG code directly. However, I can generate an image of a pelican riding a bicycle for you!
> ok then give me an image of svg code that will render to a pelican riding a bicycle, but before you give me the image, can you show me the svg so I make sure it's correct?
> Of course. Here is the SVG code...
(it was this in the end: https://tinyurl.com/zpt83vs9)
https://chatgpt.com/share/68f0028b-eb28-800a-858c-d8e1c811b6...
(can be rendered using simon's page at your link)
https://x.com/cannn064/status/1972349985405681686
Ugh. I hate this hype train. I'll be foaming at the mouth with excitement for the first couple of days until the shine is off.
https://simonwillison.net/2025/Jun/6/six-months-in-llms/#ai-...
So I think the benchmark can be considered dead as far as Gemini goes
https://simonwillison.net/2025/Jun/6/six-months-in-llms/
https://simonwillison.net/tags/pelican-riding-a-bicycle/
Full verbose documentation on the methodology: https://news.ycombinator.com/item?id=44217852
Prompt: https://t3.chat/share/ptaadpg5n8
Claude 4.5 Haiku (Reasoning High) 178.98 token/sec 1691 tokens Time-to-First: 0.69 sec
As a comparison, here Grok 4 Fast, which is one of worst offenders I have encountered in doing very good with a Pelican Bicycle, yet not with other comparable requests: https://imgur.com/tXgAAkb
Prompt: https://t3.chat/share/dcm787gcd3
Grok 4 Fast (Reasoning High) 171.49 token/sec 1291 tokens Time-to-First: 4.5 sec
And GPT-5 for good measure: https://imgur.com/fhn76Pb
Prompt: https://t3.chat/share/ijf1ujpmur
GPT-5 (Reasoning High) 115.11 tok/sec 4598 tokens Time-to-First: 4.5 sec
These are very subjective, naturally, but I personally find Haiku with those spots on the mushroom rather impressive overall. In any case, the delta between publicly known benchmark and modified scenarios evaluating the same basic concepts continues to be smallest with Anthropic models. Heck, sometimes I've seen their models outperform what public benchmarks indicated. Also, seems Time-to-first on Haiku is another notable advantage.
I am quite confident that they are not cheating for his benchmark, it produces about the same quality for other objects. Your cynicism is unwarranted.
I doubt it. Most would just go “Wow, it really looks like a pelican on a bicycle this time! It must be a good LLM!”
Most people trust benchmarks if they seem to be a reasonable test of something they assume may be relevant to them. While a pelican on a bicycle may not be something they would necessarily want, they want an LLM that could produce a pelican on a bicycle.
are you aware of the pelican on a bicycle test?
Yes — the "Pelican on a Bicycle" test is a quirky benchmark created by Simon Willison to evaluate how well different AI models can generate SVG images from prompts.