2nd only to Opus 5.5, which is perfect. https://jonclegg.github.io/pacman-bakeoff/entries/claude-opu...
Up until very recently, all models struggled with this.
All results: https://jonclegg.github.io/pacman-bakeoff/
2nd only to Opus 5.5, which is perfect. https://jonclegg.github.io/pacman-bakeoff/entries/claude-opu...
Up until very recently, all models struggled with this.
All results: https://jonclegg.github.io/pacman-bakeoff/
About half the time it made a playable game in a single short prompt. The other half of the time a few follow-up prompts were needed for refinement (eg. Things like "the blaster weapon is way too powerful, divide it's hit points by 10" or "we need a way to reconnect a player whose network dropped mid round" or "the GPS doesn't work on iOS")
this is literally faster to do it yourself
> "we need a way to reconnect a player whose network dropped mid round" or "the GPS doesn't work on iOS"
these would not.
It's literally not unless you already know exactly where it is in the code
Also even if it is I find that the extra mental switching is not worth it, that's why I even have it do basic things like updating the text in buttons these days. There is no point in using my mouse and keyboard to track down a file and then make the change when I can just use my voice to tell it what to change and then wait a few seconds.
If your working with an active context, and changes you do then need to be conversed back to the agent, and even then, it might still find it jarring and wrong.
Could it be because the model was somehow pre-trained? If we compare it with pelicans that are still not-perfect…
Fable 5.1 was pretty good. Even animating it: