It drew a better pelican riding a bicycle than Opus 4.7 did! https://simonwillison.net/2026/Apr/16/qwen-beats-opus/
It drew a better pelican riding a bicycle than Opus 4.7 did! https://simonwillison.net/2026/Apr/16/qwen-beats-opus/
I just tried this GGUF with llama.cpp in its UD Q4_K_XL version on my custom agentic oritened task consisiting of wiki exploration and automatic database building ( https://github.com/GistNoesis/Shoggoth.db/ )
I noted a nice improvement over QWen3.5 in its ability to discover new creatures in the open ended searching task, but I've not quantified it yet with numbers. It also seems faster, at around 140 token/s compared to 100 token/s , but that's maybe due to some different configuration options.
Some little difference with QWen3.5 : to avoid crashes due to lack of memory in multimodal I had to pass --no-mmproj-offload to disable the gpu offload to convert the images to tokens otherwise it would crash for high resolutions images. I also used quantized kv store by passing -ctk q8_0 -ctv q8_0 and with a ctx-size 150000 it only need 23099 MiB of device memory which means no partial RAM offloading when I use a RTX 4090.
* It's sitting on the tire, not the seat.
* Is that weird white and black thing supposed to be a beak? If so, it's sticking out of the side of its face rather than the center.
* The wheel spokes are bizarre.
* One of the flamingo's legs doesn't extend to the pedal.
* If you look closely at the sunglasses, they're semi-transparent, and the flamingo only has one eye! Or the other eye is just on a different part of its face, which means the sunglasses aren't positioned correctly. Or the other eye isn't.
* (subjective) The sunglasses and bowtie are cute, but you didn't ask for them, so I'd actually dock points for that.
* (subjective) I guess flamingos have multiple tail feathers, but it looks kinda odd as drawn.
In contrast, Opus's flamingo isn't as detailed or fancy, but more or less all of it looks correct.
https://files.catbox.moe/r3oru2.png
- My Qwen 3.6 result had sun and cloud in sky, similar to the second Opus 4.7 result in Simon's post.
- My Qwen 3.6 result had no grass (except as a green line), but all three results in Simon's post had grass (thick).
- My Qwen 3.6 result had visible "tailing air motion" like Simon's Qwen 3.6 result.
- My Qwen 3.6 result had a "sun with halo" effect that none of Simon's results had.
But, I know, it's more about the pelican and the bicycle.
I can't comment that flamingo.
https://simonwillison.net/2025/Nov/13/training-for-pelicans-...
This reminds me of Pictionary. [0] Some people are good and some are really bad.
I am really bad a remembering how items look in my head and fail at drawing in Pictionary. My drawing skills are tied to being able to copy what I see.
"Make a single-page HTML file using threejs from a CDN. Render a scene of a flying dinosaur orbiting a planet. There are clouds with thunder and lightning, and the background is a beautiful starscape with twinkling stars and a colorful nebula"
This allows me to evaluate several factors across models. It is novel and creative. I generally run it multiple times, though now that I have shared it here, I will come up with new scenes personally to evaluate.
I also consider how well it one shots, errors generated, response to errors being corrected, and velocity of iteration to improvement.
Generally speaking, Claude Sonnet has done the best, Qwen3.5 122B does second, and I have nice results from Qwen3.5 35B.
ChatGPT does not do well. It can complete the task without errors but the creativity is atrocious.
155W PSU seems to be unified with M4 Pro model, plus there's reserve for peripherals (~55W for 5 USB/Thunderbolt ports).
Apple lists 65W for base M4 Mac itself: https://support.apple.com/en-am/103253
Notebookcheck found same number: https://www.notebookcheck.net/Apple-Mac-Mini-M4-review-Small...
I thought that's exactly what they are?
We know that LLMs build complex internal representations of language, logic, and concepts rather than just shallow word-counting.
If you deny that then you probably have an elementary understanding of how they work. Not even Chomsky denies that. The real argument imo is whether those internal representations constitute an actual "understanding" of the world or just flatten out to something much less interesting.
Actualy most statistical models can "hallucinate", specifically those that are capable of interpolation.
I have witnessed this for example in Gaussian Processes. In my own scientific work.
Even the standard introductory exercise artificial neural networks, handwritten digit recognition, already shows deeper understanding. These simple networks take in raw pixels and somewhere in the many layers recognize "curves" and "edges" and then "circles" and "boxes" and whatnot and eventually "digits".
I think there's a genuine debate about whether or not this is a form of intelligence. I think the oversimplified argument of them just being stochastic sentence machines mostly comes from people who don't understand how they work. But I also think there's a much more nuanced version of this argument offered by people like Chomsky that should be taken seriously
Any specifics? That doesn't say anything about them not being sentence generators. And it's pretty well known that the LLMs constantly spew out fantastically grammatically correct sentences that have no logic to them whatsoever.
> These simple networks take in raw pixels and somewhere in the many layers recognize "curves" and "edges" and then "circles" and "boxes" and whatnot and eventually "digits".
That sounds like a version of anthropomorphizing. It is my understanding that it is a completely open problem as to what neural networks are actually doing in their internal, deep layers.
> I think the oversimplified argument of them just being stochastic sentence machines mostly comes from people who don't understand how they work.
I mean, that's effectively a logical fallacy, so it's not a strong argument.
Tthe right one looks much better, plus adding sunglasses without prompting is not that great. Hopefully it won't add some backdoor to the generated code without asking. ;)
GLM-5.1 added a sparkling earring to a north Virginia opossum the other day and I was delighted: https://simonwillison.net/2026/Apr/7/glm-51/
If we want to get nitty gritty about the details of a joke, a flamingo probably couldn't physically sit on a unicycle's seat and also reach the pedals anyways.
Stylized gradients on the flamingo
Flowers
Ground/grass has a stylized look and feel
...despite a miss along the Y-axis where it's below the seat, couple oddly organized tail feathers, spokes, the composition overall is much closer to a production quality entity
Opus 4.7 looks like 20 seconds in MS paint.
Qwen3.6 looks incomplete due to the sitting position, but like a WIP I could see on a designer coworkers screen if I walk up and interrupt them. Click and drag it up, adjust tail feathers and spokes, you're there or much closer, to a usable output
Simon, any ideas?
https://ibb.co/FLc6kggm (tried here temperature 0.7 instead of pure defaults)
Thinking mode for general tasks: temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0
Thinking mode for precise coding tasks (e.g. WebDev): temperature=0.6, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=0.0, repetition_penalty=1.0
Instruct (or non-thinking) mode for general tasks: temperature=0.7, top_p=0.8, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0
Instruct (or non-thinking) mode for reasoning tasks: temperature=1.0, top_p=0.95, top_k=20, min_p=0.0, presence_penalty=1.5, repetition_penalty=1.0
(Please note that the support for sampling parameters varies according to inference frameworks.)I'm impressed about the reach of your blog, and I'm hoping to get into blogging similar things. I currently have a lot on my backlog to blog about.
In short, keep up the good work with an interesting blog!
Is the 20.9GB GGUF version better or negligible in comparison?
I get some really amusing 'reflective' responses, but I think it needs a bit more cooking. Maybe I'll try another variant.