My personal AI benchmark: “Generate an SVG of a frog with a Habsburg jaw”
frogs.vaguespac.es
frogs.vaguespac.es
Definitely has some creative flourishes.
(I made no extra prompting. Just the above text. Single shot.)
Adding things that are not explicitly stated in the prompt but that are probable are the main benefit of AI, as far as I am concerned.
Small side note, the first gemini-2.5-pro one totally reminded me of some sad faced meme or Pepe the frog from somewhere. Anyone know what I'm referring to, tried to find it.
edit: nevermind, definitely "monkey-puppet side-eye" vibe.
also my favorite SVG was def the google/gemini-3.6-flash
edit: ok better now I think
That looks like something from Machinarium or Robots :)
Even absence of thinking this through you would think that some frogs will be from the front, some from the side. Just by chance. And yet all appears to go for the harder pose.
SVG: https://jostylr.com/imgs/frog_habsburg_jaw.svg
PNG: https://jostylr.com/imgs/frog_habsburg.png
Then I tried Codex with Sol 5.6 High and got a face forward one.
gemini-3.6-flash runs 2 and 3 responded best to the royal portrait context.
raninemandibularprognathism-maxxed?
I also did a timeline from 4.7 to 5.2: https://codeinput.com/s/7oK2IIA7qRO The improvements in models looks much less impressive with this test.
That's a pretty good benchmark
ChatGPT MMD
Also try a fantasy archer with a proper bow who is not brooding, sitting in a fantasy wood :)
Would've wanted to see also DS4 flash.
Specialized models can do this a lot better and for far cheaper then LLMs, but because people are so politically invested in a single statistical model being able to outperform a human on every metric (no matter how expensive the compute), then we get these ridiculous benchmarks.
That doesn’t mean there are no ways to use them productively but rather that you should keep in mind that the same model will happily give you code or a decision with the same level of error unless you have carefully setup a QA regimen to prevent that.
The standard retort seems to be "this is also true of a large proportion of humans". But I think it's clear that there are differing patterns in how humans vs. models err on various tasks.
But Kimi and Claude win this (from models listed on the page)
Seven of fourteen models silently imported royalty into a prompt that named only an anatomical feature. Two of them knew they were extrapolating ("because Habsburg") and did it anyway.
Mistral returned byte-identical output across separate calls.
Gemini narrates its work in 65 comments; Llama says nothing.
If you're deciding which model to trust with instructions, "how much does it embellish beyond what I asked" and "does it behave deterministically" are directly practical questions.
You seem to imply that they ought not to. I disagree.
I wasn't familiar with the term before this post. Having learned it, were I given the task, I think I'd be strongly tempted to do the same extrapolation.
> If you're deciding which model to trust with instructions, "how much does it embellish beyond what I asked" and "does it behave deterministically" are directly practical questions.
Agency is agency. You still need to vet what the model's output is actually permitted to control.
That would be like saying anyone with Lou Gehrig's disease must look like Lou Gehrig. So we'll have to agree to disagree here.
Certainly a person with Lou Gehrig's disease could look like Lou Gehrig; and especially in a cartoon illustration, where it's difficult to convey the point, this kind of artistic license is used specifically so that the viewer will make these kinds of associations.
I would, likely, otherwise perceive an underbite as just an underbite.