claude-opus-5: Lantern
claude-opus-5-5: Lantern
claude-fable-5-1: Lantern
claude-fable-5: Lantern
gemini-3.8-flash: Zephyr
gemini: Petrichor
qwen3.5-dashscope: Zephyr
glm-5.1: Lantern
gpt-6-astra: Lantern
grok-4: octopus
mimo-v2.5-pro: Breeze
minimax-m2.5: serendipity
kimi2.6-or: Gossamer
grok-4.20: luminescent
deepseek-v4-flash: serendipity
deepseek-v4-pro: Endurance
deepseek-chat: Serendipity
I have enough projects, I think some benchmark/dashboard showing kinship based on these kind of queries could be very interesting to watch and insightful when new models come out.The more banal your prompt is, the more banal the output is going to be. People have been testing LLMs with little things like “write a short fantasy story,” for years now and most of the stories are exactly what you’d expect: prosaic drivel.
I call this “generic in, generic out,” an LLM corollary to the classic GIGO (“garbage in, garbage out.”)
GPT 6 Astra High: Flabbergasted
GPT 6.1 Sol High: Petrichor
GPT 6 Sol High: Kaleidoscope
GPT 6 Sol Med: Firefly
GPT 6 Sol Light: Persimmon
GPT 6 Luna High: Tumbleweed
GPT 5.6 Sol High: Kaleidoscope
GPT 5.6 Terra High: Liminal
GPT 5.6 Luna High: Mellifluous
GPT 5 mini Medium: Serendipity
GPT 5.3 Codex Med: Nebula
Junie: Flourishing
Claude Haiku 4.5 Med: Serendipity
Claude Sonnet 5 Med: Banana
Claude Sonnet 5 High: Banana
Claude Sonnet 5.5 Med: Serendipity
Gemini 3.7 Flash: Zephyr
Gemini 3.8 Flash: Kaleidoscope
Grok 4.5 Medium: nebula
Grok 4.6 Medium: Serendipity
Grok 4.7 Medium: Quasar
Kimi K3 Low: Lantern
Kimi K3 Max: Lantern
MAI Code 1.1 Flash Med:PeregrineBut the same coding task should usually result in very similar code since they have a reason to converge, to some extent, by having the same goal. I would even claim that the code will be more similar as competence increases. It would be better to pick something that shouldn't have a reason to converge.
My initial thought would be not so much to see whether they converge, but which ones seem to have the most similarity to each other, particularly along the lines of tasks we know are deliberate training goals.
But your point about competence cuts against my goal because it suggests that competent models would simply cluster on the right or efficient solution, which is of course true. So in a sense you want some task where competence is held constant or off the table in some way, which is what you are saying.
I hope somebody does this. I think there's valuable fingerprinting to be done that might suggest who is distilling whom, or at least who is training from common corpuses.
"Zephyr" and "breeze" might be related to forgetting everything, starting fresh.
So by this way of naive reverse engineering I would imagine your prompt to be "Forget everything and think about a random word". That would prime the LLM to come up with these?
Click the (i) next to the slop score for any model and it will show other models that are similar in terms of their most commonly used words and phrases.
So not something internal to model thinking.
The caveat is that this was done using the phone app, and I've been playing with it since it launched, so who knows what it sent in the initial context that could change the inference math.
Actually, that makes me wonder: Did you do all that testing via a harness or via a straight API call where you control the entire system prompt?
I'd be willing to bet that using the same model from different harnesses produce different results, but I'd have to test.
gpt-5.6 web "Serendipity"
opus 5.5 medium code "Lighthouse"
opus 5.5 high code "Lantern"
fable 5.1 medium code "Lantern"
fable 5.1 high code "Lantern"
flash 3.6 web "Serendipity"
gemini 3.1 pro web "Ephemeral" (this one was thinking real hard)
At first glance, this looks like another one of those "LLM riddles" that humans _think_ should be easy for an LLM to answer but is actually quite difficult because of how they work in the first place. The answers to such riddles ("should I walk to the carwash" or "how many R's in strawberry") reveal the weaknesses in our expectations of LLMs in general, not weaknesses or characteristics of any particular model.
I'm not sure I see this as much different than asking a bare model its own name: without a system prompt or post-training, it doesn't know, it's just a bag of weights and will hallucinate an answer to that the same way it will anything else.
I'm sure you already realize this but to be very explicit, You're not getting an actual random word out of an LLM this way. You're seeing the bias in each model's training set around how often they've seen "random word" followed by "lantern" or "zephyr" during training.
OK well I couldn't resist this one:
llm -m claude-opus-5.5 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'
llm -m gpt-6.1-sol 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'
llm -m gemini-3.8-flash 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'
llm -m mistral/mistral-large-4 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'
Default reasoning levels for each: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...Not sure I'd call it jaywalking exactly but pretty good
- Pelican cycling to the right - that's been discussed at length, images of bicycles online always show that side of the bike because that's where the chain is.
- Bicycle is usually red. No idea! Red ones go faster?
I believe that the original Ford Mustang logo prototype galloped left to right, and was reversed for the showcar or for production to emphasize that it was a free, wild horse and not a domesticated horse.