Show HN: Pac-Bench – How well can models one-shot a Pac-Man game?
jonclegg.github.io
“Create a Pac-Man game in a single HTML page”
Each model gets one shot — no follow-up prompts or fixes.
jonclegg.github.io
“Create a Pac-Man game in a single HTML page”
Each model gets one shot — no follow-up prompts or fixes.
No-one ended up with a complete clone but I loved how we all ended up focusing on different things, like pixel-perfect graphics versus accuracy in gameplay, and how we all brought our existing skills to the challenge despite not really knowing what we were doing.
I expect if we had AI models available it would have ruined the pleasure of figuring it out for ourselves. I feel kind of sad for the next generation of developers who won't have that experience.
One-shotting something like Pac-Man doesn't prove much. At the end of the day, one-shot fidelity is going to scale more or less linearly with model size/world knowledge. Why wouldn't it?
The fact that these are at all playable and 100x my programming skill level is pretty depressing from a certain pov. The ThreeJS dude posted ab how demotivated he was to continue his work, and while I was never a dev that did much with webgl, I commiserate.
It seems like a lot of the programming is cookie cutter type stuff (where these ai programs shine) and should be easier.
They can be amazing, but promoting “make a pac man game” seems like an exercise in which model has the best training code that was Pac-Man.
That was nodejs, and it turns out people unwilling and or unable to improve things through systematic skill refinement just degrades into chaos.
Successful ecosystems are just pedantically lame enough to keep silly folks from going YOLO, but empowering enough to allow people to still have fun.
>which model has the best training code that was Pac-Man.
You mean which model is more cautious about copyright bleed-through of the $9Tn in FOSS and user code they misappropriated though isomorphic plagiarism. =3
Nodejs wasn't easier for people. JS is a horrible language where simple stuff like comparisons, array access, and member access are broken.
I personally have been working on a Three.js project with Opus 5 and 5.5 that I never would have continued with had I needed to dive into documentation by hand.
Seeing immediate results is incredibly motivating.
Whilst it is the Astra aesthetic, the expectation of any three.js game will be that it is all surface and no detail, whether someone has put the effort in or not.
Mrdoob? If so, can you show a link?
Remember when people considered you a genius for prompting with "You are a skilled writer....".
Buffered controls, different pathfinding AI for each ghost, level transitions, etc.
Gpt-5.6 sol high technically completed the assignment as well, but the gaps within the pipes used to construct the maze, the lack of pause when eating a ghost, etc made it feel far less polished.
Make a pacman game where pacman can always eat ghosts, but ghosts drop pellets.
Clearly some new RLAAS/dataset/env is being used for this now (it doesn't even seem that complicated, you have one LLM judge whether gameplay is recognizable as the original game or not and another trying to implement a logically/semantically identical version of the game). It's why the performance improvement on this workload has been so dramatic.
Everything is going to go from 0->1 on this benchmark in short order because of that.
It has the advantage of taking good screenshots and letting me know if it’s decent within five seconds of playing.
These results get better and better, but the half-baked pac mans games and pelicans are just funny.
You cannot prompt that level of brokenness. Same for those slop-posters you see everywhere.