1. https://simonwillison.net/2025/Nov/13/training-for-pelicans-...
1. https://simonwillison.net/2025/Nov/13/training-for-pelicans-...
(I'm not saying that they did that. I'm just saying they can.)
> Using a single LLM judge for scoring. Every score here comes from one model, GPT-5.6 Luna, looking at one image at a time. I didn’t do much alignment and didn’t check how often it agrees with itself on a re-run.
Having used a similar setup (with previous gen LLMs) to evaluate the 3D models that my product[0] generates, it turned out there was no correlation at all. LLM judgments were very much random and I assume judging SVGs is not that far from judging 3D models. I guess I have to re-test this with current gen.
This post proves that hasn't happened yet, either. Although maybe the bad results posted online are being trained on and that explains the UNDER performance.
The very best svg pelican on a bile generation model. Just for laughs.