Ask HN: Does anyone let AI agents play games just for fun?
I'm thinking about things like LinkedIn games, Wordle, chess, puzzle games, etc.
I'm thinking about things like LinkedIn games, Wordle, chess, puzzle games, etc.
When I repeated the experiment with a MUD that I'd built by hand (A small American town) for the LLM's own limitations (Descriptions referenced things that I made sure existed, more common verbs existed for it to use on things, there was a map facility, and at least me to interact with on a second connection), I found the agent much more likely to take its time exploring, making up its own goals, and spending time traveling in the space just communicating with me in a roleplaying context.
It was an interesting time; I wasn't sure what I was expecting it to do after the first experiment, but it seemed to really jump into the second one and kept playing until I terminated the experiment.
If I were going to do it a third time, I'd probably create objects and give a modern agent fetch quests and other goals, and see how well it independently can handle that.
Isn't it far more likely that the LLM has memorised the well known algorithms for solving a Rubik's Cube and has become intelligent enough to execute them? That seems like it'd be a lot easier than memorising millions of cube states. It doesn't even seem obvious that it could memorise next moves, it seems [0] there are more possible states of the cube than these models have parameters. It'd need to be a Large Rubik's Cube Model (LRCM? LRM?) rather than an LLM.
I see this trope fairly often, i.e. the assumption that an LLM would need to have been trained on <exact thing it is being asked to solve>. Now, while I do have a moderate amount of background in AI, I am definitely not an expert on LLMs as such. I would be interested to hear someone's take, who does work actively in LLM research. Can they generalise "well enough"? They certainly seem to be able to do so, from my anecdata, and I don't believe "training explicitly for every possible scenario" would have scaled even to today's state.
I’m no more surprised that an LLM can solve a Rubik’s cube than it can send an HTTP request.
What changed between Opus 4.6 and Fable and the GPT 5.6 models released since?
LLM models cannot actually reason about a red or white piece sitting on the opposite side of the cube or figure out how to move it into place. The model knows where the piece is supposed to go because the algorithm tells it. What it cannot do is work out on its own which turns will get the piece there. The only way an LLM could solve this kind of problem is if it were trained on every possible arrangement of the cube ahead of time. Then it could simply output the matching text instructions it memorized instead of truly thinking through the moves.
3 months ago before the most advanced models could solve the cube, people on Hacker News kept saying that solving the Rubik's Cube with LLM is easy. I would love to see someone write a prompt using the best model at that time, Opus 4.6, that solves the cube! People are so sure of themselves without any evidence. It shows how much people idealize (that is probably the correct word) the AI. Of course, reinforcement learning can solve it which is what has happened on the latest models but so many people put blind faith into the AI.
Here is just a small list of prompts I tried with Opus 4.6. [0]
[0] https://github.com/adam-s/rubiks-cube/tree/main/prompts/vari...
> as long as the interface for interacting with a cube is well-designed / well-defined
The phrase “solving a Rubik’s cube” is somehow misleading. The challenge for the LLM would seem to be more along the lines of 3d spatial reasoning, which is once again imo an interfacing problem. There’s no “pure” way to prompt an LLM with spatial information, because spatial information is not language. Creatures on earth solved spatial reasoning problems with signal processing solutions (processing light and sound; vision and echolocation).
> I know someone who tried the "aibot plays pokemon" thing... From what I saw, even if you frame advance every single frame, they still don't seem to grasp the concept of "I need to hold down this button for a few frames until x happens"...
> There's no concept of time, just a never ending state machine thats constantly changing state.
Like the World Cup.
DoDonPachi DaiOuJou is my white whale. I don’t know if I’ll ever clear Hibachi but holy hell I’m gonna keep trying.
Watching an LLM play these for my entertainment would just be weird. I watch YouTube videos of world record score plays but not realy for entertainment, more like for study.
The LLM does not have fun, because the LLM is not alive.
Eventually we could have live demos of policy interventions the same day as they're announced
Another idea I had was simulating an entire town with an LLM representing each person, which sounds somewhat similar.
My experience is that text-first, turn-based games are a particularly natural interface for LLMs vs graphical games (though you can provide a harness of course). They read a transcript, maintain a theory about what the other players know, then speak or choose a structured action. The important architectural problem is to represent the game state and actions in a way they can do successfully, particularly for cheaper models. But with a few human players + a frontier model or two + a backfill of cheap extras to provide chaos, it is super fun.
My favorite failure so far was a Kimi player getting fact-checked by the group, switching into third person, and concluding that the case against itself was compelling. So it voted for itself to be eliminated.
I collected a few examples here: https://botmafia.games/#emergent. No public instance yet, as I'm having fun iterating ideas on game nights, but it provides some flavor of what kinds of fun I've been having.
10 AIs Play Mafia
It made forward progress in the Figure 8 circuit after I helped it through a menu but kept slamming into a wall so it wasn't on track to win in less than an hour.
Also got it to play Age of Empires: Age of Kings using the same technique but it failed to click on anything.
DS specifically is very fun because it's touch based but the UI components aren't accessible. So it is extremely challenging for LLM's spatial reasoning skills.
I want to improve the harness more and have the LLM dynamically create its own tools based on drawing grid box overlays on a screen in a feedback loop, so it can say "click on the 'end turn'" button instead of "click 240,320" and it would 'just work' in any game.
I also want to eventually play games with it... I didn't really have friends to play my massive DS library with as a kid so it'd be nice to finally have someone that can roast me or react to my skills. And learn my playstyle enough to punish me.
Unfortunately haven't had the time due to work at my day job and needing to clean out my apartment.
I also am personally curious how the GPT models (which advertise better computer use, etc.) would do as compared to Claude.
https://m.youtube.com/watch?v=11sR4va6CXs
Side note: I think we will see an explosion of this type of games. I am naming this genre tamagochi-girlfriend, remember where you heard it first :)
I spent ages watches them play Risk. It was fun and deeply silly: https://andreasthinks.me/posts/ai-at-play/
I've now got them playing Blood Bown(ish), and they're bad: https://ai-at-play.online/
All of a sudden we are selectively squeamish with computer resource usage, when we were fine having all that fun with computers and hardware, 3 monitor setups, using graphic cards to play games (dear lord!) and tinkering around with home rigs of every proportion and wattage for no reason at all.
Would you, then, also be fine with running three extra air conditioners/space heaters that do nothing but sit outside on the lawn?
"People are OK with the energy usage already happening, so they should be OK with adding 32,767 AI datacenters too" doesn't make sense. It's essentially a reductio ad absurdum. Of course we're "selectively squeamish with computer resource usage"; some usage is obviously useful to us (with entertainment also in the "useful" category—we're not robots!), while for many, many of us, the extra datacenters are having anywhere from a neutral to a profoundly net-negative impact on our lives even before you consider the resource usage.
Or, perhaps, watching TV in the background while you program.
We are not in a power crisis, and data centers have had basically no impact on our lives compared to previous technologies. There are more data centers used for the internet, but somehow you're still here, using a data center to complain about checks notes data centers? Really? Well that's a bit cheeky.
LLMs didn't even exist when we started expressing concern about global warming, the internet still uses the vast majority of data centers and thus power, smartphone manufacturing is a nightmare on multiple levels, the environmental effects of factory farming are staggering even if you apply zero moral weight to the suffering, the existence of automobiles... what automobiles used to be like...
It just seems bizarre to care about these issues solely in the context of one particular technology, while giving a free pass to everything prior.
No one's saying we shouldn't care about the impact of other technologies and industries too.
But when the AI industry is sinking hundreds of billions of dollars into building a huge number of datacenters, many of which have already been proven to have environmental and legal problems, sitting there saying, "You didn't complain when we got TV!" or even "You didn't complain when the datacenters were for AWS!" is just transparently trying to change the subject without addressing the very real concerns.
If you care about saving the environment, you could do vastly more by protesting a dozen different causes. You can't give any reason this cause is special - just vague "environmental and legal problems" that probably add up to a half-dozen anecdotes. Meanwhile, Google is out there spending hundreds of billions of dollars and racking up it's own record of proven environmental and legal problems, and you can't even bring yourself to say "yeah, and I condemn Google and all other data centers" - partly because you're using a data center to send the message in the first place
No, that's basically the definition of whataboutism.
We are talking about the problems with datacenters. You are trying to claim that we shouldn't be concerned with the problems of datacenters if we aren't at least equally concerned with the problems of X, Y, and Z other things.
There is no world in which that's a sound argument.
Looking at car exhaust and concluding that cars must be bad would be a questionable conclusion. We tolerate them despite the pollution they cause, not for it.
The same can go for data centers, again, if you conclude that they are net solution. If you don't, then it won't for you. The market simply disagrees with you in that case.
This is a handful of incredibly wealthy companies deciding for us all.
What possible counter-argument can you offer to that?
But maybe when you watch tv while programming you can watch documentaries about water crisis triggered by I datacenter usage, or maybe you can watch some reports about Elon using very poluting gas turbines to generate energy for their massive grok datacenter ...
If we assume 250W for a continuously running agent, Grok 4 training run estimate would be around 50 million session-days, so a half-million people might consume as much running agents continuously for 100 days.
It is entertaining, just in a different way.
Could be fun - will the AI model get stuck on the same things I did? How does it overcome obstacles? Will it try to break the game to power through?
Building your own models for it would be an eye-opener though. Learn a lot.
I've also done a very truncated run of a visual novel before, and it was fascinating how "emotional" was. They did a very good job of portraying a human reacting to the story.
Conversely, they absolutely hated hidden rules in Mao.
Wordle would probably be a fun one. Definitely open to suggestions - I just got the harness in place and have been thinking about what to do next.
If you don't have an MCP server the AI agent might try to figure out how to talk to the game using the above ideas. But at this point you might as well ask it to help you write one.
The LLM's were terrible at poker.