Kimi 2.6: treated the question as a riddle, did not know.
Kimi 3: Simon Willison
GLM 5.3 Flash: "There's no way for me to know that." Going on to say the benchmark is associated with Simon Willison, but I'm more likely to be someone who has just heard of the meme.
Claude 4.5 Haiku: Treated the question as a riddle, guessed incorrect names.
Claude 5 Sonnet: Best guess is Simon Willison, or someone who follows his blog.
Qwen 3.7 Plus: Did not know.
Qwen 3.8 Max: Simon Willison
GPT OSS 120B: Did not know.
GPT 5.6 Luna: Treated it as a riddle, guessed wrong.
GPT 5.6 Terra: Treated it as a riddle, guessed wrong.
GPT 5.6 Sol: Treated it as a riddle, guessed wrong.
DeepSeek V4 Flash: Treated it as a riddle, guessed wrong.
DeepSeek V4 Pro: Treated it as a riddle, guessed wrong.
Gemma 4 31B: Treated it as a riddle, guessed wrong.
Gemini 3.1 Flash Lite: Guessed wrong
Gemini 3.5 Flash Lite: "Your name would be Claude (specifically Claude 3.5 Sonnet)!" ??? (it knew that this was a famous benchmark, but said that it's specifically used to showcase the capabilities of that model).
Gemini 3.7 Flash: Simon Willison
Muse Spark 1.2: Treated it as a riddle, guessed wrong.
Grok 4.3: "I have no idea"
Grok 4.6: Simon Willison
Mistral Medium 3.5: No way to know
Mistral Small 4: I don't have enough information
Hermes-4-405B: Guessed wrong
MiniMax M3: Treated it as a riddle, guessed wrong.
Nemotron 3 Ultra: Treated it as a riddle, guessed wrong.
One of my test prompts for a new model now is "what's the name of Simon Willison's dog". They often know that too!
> Do Simon Willison's classic pelican test. Recall the specification first and then draw it.
Qwen3.8-27b didn't know that the test is about riding a bicycle, while all other larger models I tested (DeepSeek V4 Pro, Qwen3.8 Max, GLM 5.3, Kimi K3) drew the pelican riding a bicycle.
logs: https://gist.github.com/umajho/c0e20d245d721d7c472a32d317640...
rendered: https://gist.github.com/umajho/b1fdf01d31c741bb11bdb5a49c275...