Instead, it feels like a more appropriate benchmark for the original purpose would be to come up with new, novel problems each time, and compare across all models (including previous ones).
Throwing my hat in the ring: Generate a pelican shaped crossword where all the clues are related to bicycles.
(Haiku 4.5: https://imgur.com/a/N112Nxo, I'm trying some others but it's very slow! Opus has been at it for about 20 minutes.)
(don’t tell my boss.)
This makes me think AI companies are using chat history to train the next model.
It may be that AI performs better on such code.
reminds me of this Key and Peele skit
AI image generation suffers from this more generally. You can generate pictures of pelicans, sure. Newer models clearly generate images with more pelican-ness than before. But all of it is still uglier than sin. Drawing things accurately is one thing, making results that someone might actually want to use (without embarrassing themselves) is something else.
As for conventional diffusion-model stuff, I happen to think there are some pieces of AI art that still look really good even knowing they're AI.
If you stick to the benchpress, it's just "benchmaxxing".
Minimizing peaks is probably not a good strategy generally. Even low peaks may have some benefit ("if you know your problem is in this domain and performance is critical, this tool offers a 3% advantage").
Minimizing troughs may be more attractive. People generally seem to react more strongly to negatives and competitors can devise benchmarks which emphasize one's troughs.
While a universal expert would be convenient and broad knowledge aids some forms of creativity, specialization has substantial advantages.
Trading max performance (primary metric of concern) for some improvement in a secondary metric often makes sense and that could be an interpretation of "minmaxxing", reducing the over-emphasis of a single metric which would otherwise be maximized.
There seems to be a very strong correlation between models that are good at SVG and models that are good at 3D CAD.
Anecdote I know, but there does seem to be generalization going on here.
I haven't really tested Opus 4.8, but 4.7 wasn't nearly as good as ChatGPT 5.5.
It could well be that the exact domain matters more than the bigger picture concepts, like 3d. One of the ever fewer reminders that this tech is still just fundamentally a token prediction algorithm.
Mostly I have been using GPT 5.5 and now 5.6
That is notable because I do almost exclusively use Claude for coding.
...which.. hmm I dunno if they are same or not
That said, I think this would correlate relatively little with general programming ability. They're not unrelated, of course, but being able to generate code that paints an accurate + esthetically pleasing image is quite different from generating code that achieves a non-spatial goal.
Why do LLMs need to be able to do this as well, but worse, slower and more expensive?
LLMs do a great job because they understand both code and SVGs well.
Edit:
An example for a synth I'm buulding: https://imgur.com/a/U694Ek7
The irony of having to post it as a PNG isn't lost on me...
Why on earth would you let an LLM do this when it would take you 10 min to do this in Figma or Inkscape or even just Word
I haven't read the code.
Why would I do it in a slower, more difficult way for something that's going to be outdated in 2 hours?
This design is trivial. I admit it would be hard to achieve in Word (or at least for me because I don’t know how to make a diagram in Word any more) but Figma and Inkascape are made to do these things, and have optimized UI for that (personally I would have just used mermaid though).
I think you may have lost your faith in human capabilities just a little bit if you think drawing stuff like this takes any time or effort at all. Compared to designing and architecturing the system, drawing the diagram is trivial. Now I know that my parent did neither but it seems like they vibe-coded the whole thing. I’m sure they will end up with a fun little toy from the whole endeavor they can play with for 2 weeks before abandoning. Maybe the author will even feel bad about the carbon footprint of this whole exercise and buy some carbon offsets to make up for it.
That's exactly the idea - except it's more like 2 hours before I prototype the next version.
> Maybe the author will even feel bad about the carbon footprint of this whole exercise and buy some carbon offsets to make up for it.
The passive aggressiveness of this is perfectly weighted and admirably phrased.
Where I'm from data centers help the renewable mix by subsidizing transmission from other geographic zones. I'm actually improving the environment by using it.
[1]https://ourworldindata.org/how-much-energy-do-data-centers-a...