'A lot worse than expected': AI Pac-Man clones, reviewed
theguardian.com
theguardian.com
But once you get past the initial wow factor it’s hard to move forward with the refinement. Recently I tried to use multiple LLM tools to implement a solution to a problem that I had solved before. I spent hours prompting and reprompting trying to tell it exactly what to do, but every time I’d pin down one fix it would break something else by gravitating back to a different solution that didn’t work. Finally I re-read the tutorial in the library’s documentation and realized that it was just pulling my code back toward the documentation example at every opportunity. My attempts to deviate and implement something more complex kept breaking it.
I guess my question is if you had a good solution you'd done before, and also a good understanding of what kinda changes it needs, why bother with hours of AI prompting?
Unless all the promises and hype are false...
Funny thing is it reminds me of when I'd first start coding. You write too much, compile, wait, compiler error. Read error, fix, add print statements, repeat. There's that down time with the compiler that made it feel not as bad but also make me lose concentration.
Similar to when I started coding it also results in things that technically work but are clearly spaghetti held together with duct tape. But at last that process resulted in me becoming better
Where I have found vibe coding as an approach really shine is if I need to write some sort of quick utility to get a task done. Something that might take an hour or more to slap together to solve some menial task that I need to do on a bunch of files. Here I can definitely throw it together quicker than manually and don't care if it is messy code.
Larger, more complicated apps that are meant for production are painful to try to get AI tools to build. Spend so much time prompting the AI to get the task done without breaking something else that I doubt I'm any faster than just hand coding it alongside a co-pilot
I find that it works best when used by an actual programmer who has a good idea of exactly what they want to do and how they want it done. I often find myself telling it extremely specific things like, add a switch case in this callback in this file. Add a command in this file after this other one. Create a new file in this directory that follows the convention of all the others. And so on. If you instruct it well, you can then tell it to repeat what it just did for every item in a list that is like 20 items long and you will have saved hours of development time. Very rarely does it spit out fully functional code but it's very good at saving you the time it takes to constantly repeat yourself.
(This codebase isn't that good at DRY, I try my best with things like higher-order functions but there's only so much I can do, I still need to repeat myself in many cases.)
As an example, each tool callable by the AI needs its own input JSON schema, and in its execute function needs to send a request to the client, client needs to have a callback that handles that request, etc. it is very boilerplatey, bridges across multiple implementation languages, in completely different parts of the codebase, and Claude knocks it out in like 30 seconds flat so I can focus on the parts of the implementation that it slightly fucked up, but it usually gets the boilerplate bang-on.
I still do a bit of manual post-processing but the end-to-end latency is a lot lower than if I had done all of the processing that manually.
Another relatively common pattern for me is to do a certain transformation manually once and then tell Claude Code to repeat it on all of the related tools. It does a good job at that too. I don't even have to manually tab through each file in my editor, or come up with some cursed find and replace pattern, or anything like that. It's just more time-efficient for me.
Note that I don't use an LLM to save mental effort. It's purely for saving time. People who try to use an LLM to save mental effort are usually using it wrong. You still need to know what you're doing in order to properly tell the LLM to do that.
This reminds me of the fad of creating Twitter/Reddit/Hacker News clones in 15 minutes in some new web framework back in the 2000s/2010s. And of course none of the 15 minute Twitter clones actually have the hybrid fan-out architecture that Twitter wrote after a few years so that it wouldn't failwhale all the time. I'd love to see someone vibe code their way through that.
What I don't understand is that the same people who embrace long hours, grinding, and having secret founder DNA sauce also want to completely take their hands off the wheel and vibe code their way to financial freedom.
Note that just turning on Copilot/some other assistant and having it do code suggestions is fine in the sense of surrendering far less control over architecture and correctness and is something I do at work to cut down on boilerplate. So I think there is a spectrum from micro-AI assistance to macro-AI vibe coding. For instance, you could ask the chat bot to help you implement specific parts of the app, like using Djikstra's algorithm for pathfinding for the AI, without just asking it to make the entire app.
Other models often perform better (https://web.lmarena.ai/, https://aider.chat/docs/leaderboards/) - I have yet to meet anyone who uses Grok as their primary programming assistant.
Aider can't use grok 3 with thinking yet, afaik, because xai hasn't made it available in the API.
From what I'm hearing, it and Claude 3.7 "thinking" are very similar in performance.
Also this is hilarious:
Unfortunately, my invitation to chat further over Zoom apparently means I’m after his crypto. “[Expletive] you, scammer,” he says. I ask if it would help if I sent over my Guardian credentials. “How about [Expletive] you?” comes the reply.
I recently made LLMs play Minesweeper and ALL LLMs that I tested had a pretty bad win to loose ratio. Like the only model that won more than 3 times was R1 (mind you there were 50 games).
But AI coded software for planes or pacemakers?
ChatGPT was extremely useful in getting started, especially since in the beginning I didn't feel too comfortable with Swing. But as I built the project and the logic got more complex, AI became much less useful.
Overall it took me about 30 hours to build a complete game meeting all the requirements of the project and on which I got 100%.
I didn't lean on AI as much as I could have -- point of the project was to get better at Java, but I think that even if I did, that time would have been cut down to 20 hours.
So the AI is already confused and spouting nonsense from the start. An "accurate version" which is "complete with [...] smoother gameplay"? Smoother than what? Smoother than what would be accurate? Smoother in what way?
Perhaps it's regurgitating a mangled version of a description of some improved implementation where "smoother gameplay" (compared to some reference) would make sense?