However I do agree that the results are much worse when you try to use them. They look great in screenshots and video clips which makes them perfect for content farmers.
All of the LLM generated games I’ve played have been really bad to play, though. I even tried my hand at a simple game, thinking I could iterate on it with prompts to fix some rough edges. After the initial productivity burst every change turned into a slog of tokens with one thing changing and something else breaking it. I would try to use my remaining weekly token budget across Anthropic and OpenAI to refine it at the end of every week but after a couple weeks it felt like I wouldn’t be getting anywhere without scrapping it and going back to having the LLM build it one step at a time with my careful instruction.
Which, in retrospect, is the only way I can get usable output of an LLM for anything complicated, so it’s not surprising. It’s a fun reality check project though.
On the one hand I think most of us are incredibly impressed because we know, that quick demo would have taken us months of work to build in the before times.
On the other hand the promise is a cure for cancer and the end of all work.
So when everyone is telling you “skill issue is why you can’t one shot WoW”. It’s hard to know how you’re supposed to feel about Karpathy advertising one shot custom virtual worlds but giving you slop. Incredibly impressive slop when compared to how long it would take to create it just 3 years ago, not so much compared to Elon saying - “by the end of this year, grok will create a version of the odyssey that competes with Nolan’s”
I guess this is the average story and, similarly, the average game is boring and predictable
Early generative AI at least had the virtue of relentless, unsettling weirdness, in the same way that generative art from the late 90s and early 2000s did. A handful of people made creative use of that spooky weirdness.
Now it turns out "Airspace" art.
If you took the best, most creative, human writer in the world, and for thought experiment reasons they had amnesia (to mimic AI blank context windows) specifically while you asked them for a story idea 100 times in a row, my expectation is that this human would also give you the same idea at least 80 times out of that 100.
* still better than the mean human, but even the top 0.1% of humans aren't all professional authors.
What you think is better is not what I think is better. Imagination is not storytelling. You ask the "mean human" to write, it's going to be worse than an LLM in spelling and grammar, if you get anything at all.
Most humans never do a creative writing course after school, and the longest fiction most people will write is their resume description of what their previous jobs involved, or perhaps their dating profile.
Don't mistake what you see published (or what your friends are like) for the average human: the average Hacker News comment easily above the writing grade (and creativity) of e.g. many of the one-shots stories I've seen attempted on some creative writing subreddits. "Mean" is not a high bar.
This should cause you to reconsider your beliefs leading up to it.
GPT-4 (!) has, in studies comparing multiple models with humans for creativity, beaten the mean human. In comparison, even just the mean of the top 50% of humans beat the models studied in that case (link follows), but the point is that if you think the mean humans is particularly noteworthy, you've avoided the half of the population who have the creativity of a pot noodle.
https://www.nature.com/articles/s41598-025-25157-3/figures/2
There's more studies out there with other types of creativity test, but the conclusion is basically the same: the best models beat the mean human, but are nowhere near as good the worst *publishable* human.
My guess is this is both why LLM slop happens and why it grates so hard: a significant number of bosses look at the output and think to themselves "wow so creative" because it's more creative than they themselves are; but those bosses weren't hired to be creative, they were hired to be a boss, and all the people who they hired to be creative are going "arg, no, can't you see how bad this is?"... but that is just a guess, I've not found any surveys comparing *management* creativity to LLMs, closest is e.g. this about decision making, not creativity: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5156585
I did get pip, but it's a story about trying to rescue the moon from a lake and the raccoon lives on moonberry lane
What's pip from?
So "moon" is also part of the predigested artifacts. Maybe because raccoons are nocturnal?
I guess the next question is if they can be made fun without too much additional work with a human guiding the AI.
They can't. If you think about how these things are trained it's blatantly obvious fun is an impossible metric to optimize them for
..feels like there was a black mirror episode about that though
That's oddly specific and it'd really hurt my cousins feelings :)
These conversations get frustrating since one group is saying the dancing is bad, a second group is saying it's good (for a bear), and a third group is saying we are a few years from a bear-only dancing industry.
There was a vision of the Hololens being a general purpose compute device that enabled something more like the power of a desktop without being tethered to a sitting position at a desk. You could see it if you squinted hard (incidentally, the device caused a lot of eye strain). But even at the height of the VR hype cycle, I didn't know anyone pushing the Hololens on anyone. Nobody was telling the world they were going to be using Hololenses v1 or v2 as their only computer "or get left behind."
One would hope it accelerates AR/VR - if decent studios finally get behind it.
If I see one more “one shot MMO” where you just walk around and do absolutey nothing or another menu slop idle battler or rogulike deck builder I’m going to go Postal in Minecraft.
Used to be if a game looked that good it probably had time spent on the game part too.
Whenever I think about AI games I find myself thinking about Tiny Wings, Flappy Bird and Angry Birds. Three simple, elegant games.
It is easy to see what makes Tiny Wings so completely loveable — it has a sculpted, adorable, perfect charm with a cleverly inverted game mechanic that has a calibrated level of exasperation and reward.
But why were Flappy Bird and Angry Birds, very basic games with very old game mechanics, so charming?
It seems equally impossible to imagine an AI coming up with a game with the quality of any of them, even with maximised creativity. But explaining why for Flappy Bird seems quite difficult, especially when you consider it uses some stolen visuals!
Maybe it's the smoothness of motion that makes these games understandable and LLMs seem to consistently fail at that. Ask them to do something snowboarding and they go really hard on the physics since it seems like they don't know what kind of approxmations feel good.
(It's been a while since I was in the game industry, so IDK quite how accurate this is).
These one-shot products aren't games. They're barely even demos. I don't even know what to call them. For a mature framework like Phaser to sell-out like this and create a vibecoded platform for vibecoded games is shocking.