https://simonwillison.net/2026/Apr/2/gemma-4/
The gemma-4-31b model is completely broken for me - it just spits out "---\n" no matter what prompt I feed it. I got a pelican out of it via the AI Studio API hosted model instead.
https://simonwillison.net/2026/Apr/2/gemma-4/
The gemma-4-31b model is completely broken for me - it just spits out "---\n" no matter what prompt I feed it. I got a pelican out of it via the AI Studio API hosted model instead.
My guess is that they found a bug with their implementation of the model using the weights Google released. These bugs are often difficult to track down because the only indication is that the model is worse with your implementation than with someone else's.
This particular instance was a fix to the output parsing [1] in LM Studio, described like this:
"Adds value type parsers that use <|\"|> as string delimiters instead of JSON's double quotes, and disables json-to-schema conversion for these types."
[1]: https://github.com/ggml-org/llama.cpp/pull/21326/commits/a50...
edit: formatting
https://clocks.brianmoore.com/
but static.
I tried their model and asking a few different svg of pelicans. it is INSANE.
The training no doubt contributed to their ability to (very) loosely approximate an SVG of pelican on a bicycle, though.
Frankly I'm impressed
For example, I used to get verbatim quotes and answers from copyrighted works when I used GPT-3.5. That's what clued me in to the copyright problem. Whereas, the smallest models often produced nonsense about the same topics. Because small models often produce nonsense.
You might need to do a new test each time to avoid your old ones being scraped into the training sets. Maybe a new one for each model produced after your last one. Totally unrelated to the last one, too.
Simon and YC/HN has published/boosted these gradual improvements and evaluations for quite some time now.
There is a https://simonwillison.net/robots.txt but it allows pretty much everything, AI-wise.
Comparing bicycles between LLMs doesn't really tell us much, since how do you differentiate an AI with a good model of a bicycle, but that does a poor job of drawing one with SVG, vs one that that has a much worse model but is in fact doing a great job of rendering it?!
I suppose you could say the same for the Pelican, although it does seem more reasonable to guess that most models could accurately describe the body plan of an animal even if they can't do a good job of drawing one with SVG.
No cheating and looking at pictures. Pen and paper. Do the easy bit first and draw wheels, seat, handlebars, pedals and chain. Add a stick figure riding it if that helps.
Now draw the frame.
Now google a photo of a bicycle.