Damn I hate this benchmark. SVG authoring from head without visual reference is so wrongly posed.
(I don't have access yet.)
llm "Generate an SVG of a pelican riding a bicycle" --save pelican
When a model comes out I first make sure LLM can talk to it - usually by updating the relevant plugin, but if it's on OpenRouter I can use it directly with https://github.com/simonw/llm-openrouter - sometimes I use this mechanism instead, for OpenAI-compliant API models: https://llm.datasette.io/en/stable/other-models.html#configu...Then I run something like this:
llm -m gpt-6-astra -m pelican
Then I grab the most recent log export as markdown: llm logs -cu | pbcopy
-c means most recent conversation, -u includes token usageI paste that into https://gist.github.com and then paste the resulting Gist URL into the URL tab on https://tools.simonwillison.net/markdown-svg-renderer
If the model supports multiple reasoning levels I run it once per level and put those in the same file.
I really should automate this a bit more.
Vibe coders want a model that makes them rich, without having any actual specific idea. They write a very ambiguous prompt and expect to be amazed by the result.
Very very unrealistic and wasteful.