A generalist AI agent for 3D virtual environments
deepmind.google
deepmind.google
Out of all the fields that human do professionally, sports will be one of the last ones to disappear. The fact that it is (unaugmented) humans competing is the entire point.
This is my thought/hope for what we'll expect in the coming years as AI's automation becomes more commonplace. Society's interests will start going towards activities that showcase human ability - sports, livestreaming (very much its own industry now, but mostly for socializing, art, and gaming), performance, dance, etc. Sure AI can 'do' these things, but not at the level elite performers can or with the subtle nuisances in human personalities.
Even in the future when the AI is provide everything and we are no longer able to understand it, humans will be doing human competitions, playing chess, etc... The human on human action will be only thing left, and only thing humans care about. Chess is already unwinnable, but humans still want to measure themselves against other humans.
Chess, Go, what next? Pizza delivery? Accountant Simulator? Humans are already being outclassed one feature at a time.
I still consider myself a 4/10 at best compared to my amazing peers who studied this from the start, but you have to start somewhere!
https://gist.github.com/dfarhi/66ec9d760ae0c49a5c492c9fae939...
I feel like being able to look inside the clockwork should not distract us from being amazed at how wonderful a system can be.
I suspect the more complex the game, the bigger the advantage over humans.
In more complex games, there are more switches and the current set of best switches changes faster. With more switches, it's harder to know which are the best switches because the future is less predictable. And even if we figure out the best ones, they might change before we flick them. And even if we get around to it in time, we might fat finger it and accidentally flick an adjacent switch. And our opponent never gets tired or injured.
This is why I suspect AIs have a much higher ceiling even if we limit them to half the APM pros have. Better strategy matters less, but I admit it's our only chance lol.
FWIW, I've never played Dota but I've played a lot of AoE2 and from what I know they're similar enough (but maybe someone can correct me).
I would broadly break it into things that are complex to perform (crazy APMs or accuracy), things that are complex to understand (the stack or layers in MTG), and things that are complex to predict (e.g. time-delayed abilities and the correct time to use them, like Baptiste's lamp in Overwatch).
AI have basically constant performance across a performance complexity curve, because the complexity typically derives from physical interfaces the AI doesn't use anyways. E.g. their APM is not limited by how fast their fingers can physically move.
AIs do very poorly on tasks that are complex to understand. The best Magic: The Gathering AI's I've seen are awful (though also likely far less well-funded). Best-case scenario is basically an AI who makes plays that don't make any sense, but are at least valid plays. It's a crazy difficult problem. E.g. there are various ways to make infinite mana with combinations of cards, and the AI needs to a) realize that it can use those cards in order to create infinite mana, and b) that it can do this multiple times (I.e. it can pay for a spell that costs more mana the loop generates by going through the loop multiple times). That's very hard thing to do; human players somewhat frequently don't realize when they have loops.
Add on top of that that a game of Magic can enter a state where a loop of effects becomes recursive but doesn't result in either player winning. The game is a draw, because it cannot progress anymore. Detecting these can be non-trivial, because they might involve side effects that look like someone should win (I.e. you lose a life and I gain one, then I gain 2 life, then you deal 1 damage to me, then I gain 1 life, then you deal 2 damage to me. Life totals shift around, but net to 0 by the end of the loop).
I think the AI do well at complex prediction tasks as well, by nature of their response times and access to prior information. I would expect an AI to beat humans by a wider margin the more complex the prediction gets. Humans have finite time and thus experience; the AI is going to have more "experience", and be able to recall it at a faster rate.
* I'm not sure exactly how many heroes there were at the time, it was less than the 124 there are today, but it was certainly a lot more than 10.
AlphaGo was strait up same.
AlhaStar did have some limits placed to narrow it down for the AI. But was still imperfect information, and wildly complicated.
And those were all 3+ years ago.
Games are a lower resolution representation of the 'real' world. And we haven't seen any slowing down of AI scaling up for more and more complex world views.
Eventually the 'map' will be the 'real', as real as the human brains internal map of reality.
We absolutely have. We have superhuman performance on Go, we have human expect level performance at Starcraft, and now we get human baby level performance at 3d games. The more complex the task/game the worse the AI gets relative humans it seems, I don't see how this shows the AI scaling up, to me this is all moving horizontally.
But you acknowledge that things slowed down as we moved into more complex domains? Then you agree with my comment, the person I responded to said that things didn't slow down as we moved towards more complex domains, but there is no way you can say they haven't. AlphaStar and AlphaGo quickly competed and could beat top humans, the domains they worked with after that went way slower and still can't compete with top humans.
> Each level of capability you described was science fiction stuff before they were achieved. They’re not in any way horizontal achievements.
The second statement doesn't follow from the first, moving horizontally by applying the same things to a new domain can still unlock massive capabilities that we didn't have before.
A self driving car 20 years ago could be in a drag race and do 100mph.
Now self driving goes most everywhere, but at the posted speed limit.
And DeepMind did go off and tackle increasingly complex areas in other fields.
Not sure how anyone can argue AI slowed down. Maybe advancements in one particular game slowed down, but was that because they hit a wall? or because they shifted company resources after a game demo.
I'm not sure how you are measuring time, do you think that because AlphaStar was a few years ago, that means AI advancement slowed down? Because there wasn't another breakthrough in AlphaStar? Because they didn't keep going and fully build out every race and unit?
Do you think DeepMind has been throwing resources at Star Craft and just hitting a brick wall?
It was a proof of concept, they beat some humans, and moved on to other things like Protein Folding. Was the Protein Folding not impressive enough to think AI was still advancing?
AI is continuing to advance, because for each iteration it is tackling bigger, more complex, problems.
----------
The time between each breakthrough does seem to be a down line, less time between each plateau.
Chess : Board with a lot of possible moves, but 'manageable', the AI could just calculate every move.
GO: More possible moves than atoms in the universe or something. So the AI had to use some form of 'intuition', it could no longer brute force calculate every move. (it was only few months after AlphaGO won that they turned the same engine on Chess and it 'learned' from scratch to be a Master in only a few hours.)
SC2: There are no 'moves' it is all real time movement, and most important, there as imperfect information. The AI had to scout and keep track of un-known positions, to remember and anticipate.
Dota: Honestly, I'm not sure what the big breakthrough is for Dota. But quibbling over how many 'hero's the AI had access to seems pedantic. Wasn't this years ago. This isn't a knock on AI, the Dota work was years ago. We're arguing about an AI that is 3+ years old now. I really don't think that you can say AI research slowed because a company stopped throwing money at a game demo.
Protein Folding: Hey, lets stop just focusing on games and do something to help the world.
Poker: Wasn't Poker also conquered in this time frame, in last 2 years? Showing ability to bluff?
3D Virtual Environment: Was in discussion in another thread where everyone's main argument was AI isn't 'embodied' in the 'world', doesn't 'live' in the 'world'. And boom, same day, another breakthrough covering that. Giving machine what we would call 'vision', to understand objects in the world.
------------
Now slap this into a robot, give it a gun, and tell it the world is a 3d game.
LOL.
"We haven't had a miracle in the last 6 months, oh no, AI advancement is slowing down."
Machine learning has been a thing almost since discrete electrical circuits have existed, just wildly impractical to make generalizable versions of until recently
Humans are evolved in 3D world. Our ancestors didn't decide who can have food or sex through chess competitions. The human brains have been "trained" and optimized in 3D world intensively.
With that said, we’ve got to strangle this meme. ML/AI moves forward in unpredictable fits and starts, it doesn’t follow e.g. Moore’s law in exponential formulation.
When they’re doing research and not PR, researchers talk about “performance” on “tasks”, and define those terms rigorously.
People have been trying to improve performance, as measured by some metric or metrics, on any number of tasks, since at least the 1950s.
Certain periods of time generated breakthrough after breakthrough and a bunch of “well we’ll just scale it up and it’ll be a thinking machine” sentiment amongst the lay or semi-technical public, and similar grandiosity from experts when PR and/or funding are the objective. The world we live in I guess, but not a fire we on HN should be pouring fuel on.
During other periods of time, we’ve hit the effective asymptote on the techniques thus far invented, the scaling dimensions flattened out. Then it’s all “AI was a fad, it’s hype, this is AI Winter”.
There’s no robust consensus on when these summers and winters happen, how long they last, how much performance on one task is amenable to “transfer learning” regarding another task. It seems pretty random, the constant being the PR/funding talk.
The years since AlexNet in 2011, word2vec in 2013, ResNet in 2016, Attention is All you Need in 2017, the GA on GPT-3 series just over a year ago, and countless other interesting things have been wildly fruitful, we’ve been on a hot streak.
This is generally good news! The human race has new capabilities, win! But it’s a nearly impossible claim to defend that modern attention transformers are the final word on this area of endeavor, progress since then has been substantially brute-forced via unprecedented budgets achieved through subsidy of one kind or another, and there will continue to be periods of rapid progress unlocked by key insights, and there will continue to be less explosive periods of progress, and it serves no one with a plan more noble than “cash out while the spice flows” to tee up another collapse in interest, funding, and attention by pulling a Yud: a log scale and a ruler are never the complete toolkit on forecasting novel research.
This stuff is incredibly cool stated as flat, consensus, rigorous science, it’s incredibly exciting to practitioners and laypeople alike without any breathless hyperventilation at all. The story thus far needs no grandiose embellishment to be thrilling.
But the absolute best case in terms of research we currently know about as hyped by those seeking funding would be a nightmare end-state if it landed there (it won’t, but this disaster comes in degrees): right now the off-the wall exhilarating tech demos are so expensive that the public is effectively a spectator. There’s talk of multi-trillion dollar buildouts under complete, utterly unaccountable, ethically dubious control of people who crossed the “yikes is that even legal” line some time ago.
A trillion dollars in 2024 is give or take thirty Manhattan Projects, the idea of handing that kind of scope to people who answer to no one, hold strong minority worldviews, and give the public the finger in print by calling the bribery department “OpenPhilanthropy”?
Who the fuck thinks this isn’t a dystopian horror movie outcome in an already hyper fragile world?
It’s time to squeeze the water out of these bloated, money is no object models, make them run on reasonable power budgets in the hands of John Q. Taxpayer (who along with a bunch of helpless civilians and service men and women, ultimately foots the tab when Nadella or Riyadh write blank checks one way or another), reform copyright law so that the commons isn’t vacuumed up, compressed, and copyrighted, and take a few whacks at shit like Jenson and Lisa Su being literally cousins while partitioning the market and gouging via API lock-in.
The hyper, hyper-elite stand to gain even more immunity from all scrutiny, consequence, accountability, and even bad press if “AGI” turns out to be a mere 1-3 trillion in de facto blood money away from being locked in a vault somewhere.
Literally everyone else stands to find out that slavery isn’t a strong enough word for what this would mean for them.
You went from:
Start:
"we’ve got to strangle this meme. ML/AI moves forward in unpredictable fits and starts, it doesn’t follow e.g. Moore’s law in exponential formulation."
At End:
"Who the fuck thinks this isn’t a dystopian horror movie outcome in an already hyper fragile world?"
Isn't that hype? By the end of the post you are doubling down on the over-hyped meme's.
You're exactly right that my comment veers from high-quality to low-quality linearly with character count: I had an ambient distraction burst into my office in the middle of writing it and I was over-multitasking and failed to clean up the second half within the edit window.
The second half of my comment has important signal but it's too high-noise to be a good comment, as my grandmother used to say: "A barrel of wine and a spoonful of sewage makes a barrel of sewage".
If anyone deserved a downvote it's me, please know that it was unintentional.
It is funny because as you degraded, I agreed more. The dystopian hell scape is on the way. So hard to tell where people are coming from, from posts.
Generally, I think part of problem, we are becoming desensitized to progress.
I know the word literally doesn't mean anything anymore, but Jenson and Lisa Su are first cousins once removed. Jensen's grandfather and Lisa's great-grandfather are the same person.
https://www.businessinsider.com/nvidia-jensen-huang-amd-lisa...
Their family tree according to Jean Wu, a former Taiwanese journalist : https://www.facebook.com/photo.php?fbid=350633737719840&set=...
I don't know how their family works, but in mine and most people from my neighborhood, a first-cousin once-removed is fucking family, they're blood. Not being a securities lawyer myself I'm not sure which definition, statue, or regulation would apply here, [3] seems close (and has a creepy rush-job feel about it that smells vaguely like Kushner shit of one kind or another, Feb 2020 on an accelerated basis?).
But whether this squeaks above the line of regulations and laws and whatnot getting midnight "lgtm" stamps in an election year is, I'd argue, substantially missing the point.
When I recently said:
"Now did Lisa Su decide to "concentrate on the supercomputing market with the MI300XYZ" and Jensen decided to "concentrate on AI with Hopper" independently to a degree where the market is perfectly partitioned? Who knows, I certainly don't have proof one way or the other. But if someone made a call being like "I'm thinking of focusing on X but don't really see our differentiation in Y. How's Cathy?", it wouldn't be the fucking first time." [4]
I thought at the time I was kinda pushing it with how flip that sounded, but lo and behold, I was insufficiently cynical.
So when I say that I literally don't understand why anyone is defending this trivially dubious cartel behavior complete with a 55.58% Net Profit Margin in an ostensibly competitive market both directly and indirectly subsidized by the taxpayer (TSMC isn't going to fight off the PLA with their next process node) [5], I think Leona Lansing knows that the public will burn the building down with this shit in it before they let this shit get much ickier.
[1] https://www.youtube.com/watch?v=_1kETLlGn-8
[2] https://www.quora.com/What-is-the-difference-between-a-secon...
[3] https://www.winston.com/en/blogs-and-podcasts/capital-market...
[4] https://news.ycombinator.com/item?id=39362196
[5] https://www.businessinsider.com/nvidia-ai-chip-semiconductor...
Are they though? In the real world you’re playing at many unbounded activities at the same time, with no reward counter
But it's very cool how the OpenAI matches ended up making mid players reevaluate how they used consumable regen.
The AI didn't follow "best practice" because it wasn't trained on human games, found a better way and that was quickly adopted by all, becoming the new best practice.
caveat: my Dota 2 knowledge is lacking because I haven't followed the game for about a decade now and I have essentially 0 experience with League.
One constraint to those showmatches at the time was that every heroes had their own courier, and player at that point were not accustomed to using it for "low value" travel, unlike the AI that was using it liberally.
In a later patch, the 1 courier per hero feature was added, and now pro players are much better at managing it, but at that time it was truly a heavy opportunity cost.
Also, according to this Q&A post on reddit (https://www.reddit.com/r/DotA2/comments/bf49yk) the consumable purchasing logic was scripted, not learned.
It's worth noting here that most of these comments are missing that another dimension of the game that's completely absent which heavily influences decision making of normal gameplay: communication and progression from other lanes in the game. It's almost kind of a long running joke in the game that you'd laugh if someone asked you to 1v1 mid because it usually meant you beat them technically and they're grasping for straws to show superiority, despite how segmented and different from normal gameplay it is and how useless of a skill beating someone in such a constrained environment is.
To this point, there were better manually crafted "AI" bots that could team with eachother effectively at a higher level than the average player since the original custom map in 2003-2005. The breakthrough here IMO wasn't that it was actually making any novel decisionmaking but that it was able to perform at a high level and improve by conventional ML training, which I think is a separate callout than most of the stargazing done in the comments here.
I wouldn't say profeciency 1v1 mid (specifically the even MORE watered down rules applied here that you automatically lose after only 3 deaths or the tower is taken) translates accurately to anything in the original way you play the game unless your 1v1 matchup has a similar expectation of sitting parked in the lane, and even then it translates poorly. Sacrificing a death to kill a tower and spending all your gold so your effective loss is minimized is a legitimate trade, but in this fake constructed scenario the win/loss condition is already met. You approach the two entirely differently and more importantly, more simplisticly. That doesn't even address that some hero matchups have intentional designs to be weaker earlier in the game and/or are meant to participate in fights with multiple heroes or doing secondary objects and can't assert the same posture which goes completely unaddressed by this narrow slice of gameplay.
All that buildup to say that healing potions and staying in the lane have been a tenant of normal gameplay since it's conception, and the expectation has shifted from patch to patch. What was "discovered" here is that if you don't optimize for longer term gameplay like you would for a normal game and do the most you can to optimize for a narrow slice of early skirmishes, potions have a higher cost value effectiveness. Not sure anyone beside laymen to the game thought that was a revelation.
I think you are not giving AlphaStar the correct spin.
They came back and changed it to only have the same viewport as the human, it could not see all of its units simultaneously, it had to move the cameras like a human.
BUT importantly, it NEVER had perfect information. It could only see exactly the same as the human, just at one point they were letting it see the whole map without changing the camera, but it still could not see enemy units without sending a probe.
And. Little unsure on what the argument about APM is saying. It was slowed down to match the human speed, but somehow that makes it less impressive? That is just making it more 'human-like'. Kind of like people today want to put guardrails on AI, but if it was unleashed, it beats them easily. That isn't a knock on the AI. The AI would still have to think about every move, and form a strategy. They slowed it down to human level inputs, handicapped it, to make it playable to a human. But to your point, if AI could make 400 APM and human had 400 APM (both limited to same), then that is better measure about the 'thought' behind each individual move.
I still remember watching one match where the human was winning, the AI was down, and the AI really did fight back very aggressively from a loosing position, like a human. by expanding and adapting, and it looked very scary.
I'm stunned; how would you think they are contradictory? Imagine a transportation that moves with the speed of 1000 km/h. Very impressive, right? Now imagine media everywhere say it moves with the speed of light. Wouldn't this be over-hyping?
> BUT importantly, it NEVER had perfect information. It could only see exactly the same as the human
Maybe we're speaking about different events... In the one I'm commenting on, the AI had some zoom-out, I think 2x (meaning it would see 4 times more at once). Yes it had fog of war, but a zoom out like this is a very significant advantage.
> And. Little unsure on what the argument about APM is saying. It was slowed down to match the human speed,
No it wasn't, not exactly. Imagine that you measure a human racer speed in km/minute, every minute. Then you take the highest measured "average per minute", and program AI to move with that speed at all times. Then you praise AI for its pathfinding algorithm, because using that speed, it beats the human racers.
Yes, if a human racer has to slow down, because e.g. the human is unable to avoid obstacles at maximum speed, it does make the AI being able to move faster, impressive. But few people here would be impressed by a high reflex of a computer, because we all are used to the fact computers can react much faster than humans. It is misleading, however, to allow AI to move faster, and then give it the "spin", as you say, that the AI has won because it was smart, as opposed to being fast.
BTW, I think the AI was either only using one race, or was playing only against one race. This one thing was actually mentioned in the event (once). The APM was mentioned too, I think, but the nuance I describe unfortunately wasn't mentioned.
It makes me sad, because as I said, it is a very impressive technology. But it's hard to fully appreciate something when it is so blatantly over-hyped and when you see so many people around you being mislead and praising AI for achievements that it didn't exactly accomplish.
It could only play 1 race, but I think the opposition could be different races. I think it was protoss, but it was playing terrans and zerg. There might have been 1 or 2 units that were also removed.
In the first events where it was really dominant. It could see the whole map at once, and move its units all over the map by seeing it all. So moving units on both sides of the map practically simultaneously. BUT, this was called out as just too much of an advantage, so they made another version that actually had to move the camera around the map like a human. And the second version was still able to perform.
Map wise though, for AI, I think the dealing with the fog of war and un-known/imperfect information was the big break through. Not the map size or speed. It still had to scout, and keep up with enemy movements that were hidden, and anticipate. The zoom out didn't provide that.
I'm not totally buying the APM argument (but by end of the paragraph I do). Even if a computer can move 'faster', each move must mean something, do something worthwhile. So the computer must think out its moves. I know micro in SC2 is very big deal, and speed is essential, but you do have to know what to micro. The computer having 1000+ APM was called out also, and they added a limit. By throttling the AI to what a human can do, is handicapping the AI which isn't proving that AI isn't as good, it is showing that it can be better. Or another way, in Chess or GO, there is a time limit, but nobody is throttling the AI CPU to the same speed as a human brain, like limiting its computation cycles.
So, guess in end, I do agree, throttling APM is like a real time imposing a time limit to each move, like in Chess.
For Over-Hype/Impressive point. It is difficult. Both can be true. And everyone on the internet has different thresholds for what they think is over-hype and what is impressive. AI seems overwhelmed with both sides right now. Seemingly new miracles every day, and also over-hyped companies pumping their stock by adding an AI sticker to every product.
I just say, those AlphaStar matches were like 5 years ago, and they still blow me away.
Hard to imagine what could be possible, with this latest release in this post from deepmind.
Plug into a camera, on a robot, with a gun, and tell it the world is just a 3d game.
> By throttling the AI to what a human can do[...]
I think this is far from true, a human cannot keep the APM throughout the game on the level the AI was throttled to. If the AI was throttled to the average APM in e-sports, that would be more fair, but it was throttled to the HIGHEST APM reached by a human. Again, it's not "highest average APM in a single match", it's just highest number of actions in a given minute [I don't know if it was an all-time record or just some arbitrary value inspired by some arbitrarily chosen local record; what I know it was way too high to be fair]. Furthermore, SC2 players spam unnecessary clicks to keep themselves warmed up - if playing against AI, that is limited to average APM of its opponent, was a thing, then I'd safely bet the APM of that player would decrease 3 to 5 times WITHOUT the player reducing the number of actions that are of little (but still some) significance in order to abuse this AI limitation.
Again, the event was cool, but it makes me sad the technicalities weren't communicated clearly, which made it an advertisement rather than sport IMHO.
I do get that. I agree.
Maybe top pro's keep APM at 400 over entire match, but not really, there are ebbs/flows/spamming. While AI can max out at 400 and do that the entire match, and it isn't spamming, so every move is probably meaningful.
Maybe, throttle both AI and Human both to 200? Something like that? So both are capped lower.
For real time games. This 'throttling' is tricky. For turn based games, time limits are equal. But real time, because if we are measuring AI performance, it could be unleashed and be faster than a human and win. So is slowing down the AI really allowing for measuring the AI performance?
Like in real life.
Lets say you have a robot with a gun, and a human with a gun.
They both need to draw, aim, and fire.
Would we 'slow down' the robot to match the human? That doesn't seem like the way to measure how 'good' the AI is at doing those tasks. It could be faster.
Netflix has documentary on AI. Military had AI flying F16's, and it could beat all the best human pilots. Of course, No slowing down the AI.
In real world there are physical limits. Just need someway to translate that to real time games.
Very subjectively, I'd say: limit the APM to something very low, below 50. You now change the game: it's no longer about making many decisions, it's only partially about reacting quickly, rewarding thinking through your decisions before ordering them. This would measure the intellect more than speed.
BTW there is a mode in coop mode that AFAIR makes you pay minerals for each action, so such throttling is within canon ;)
Of course, being Bronze, with a 50 APM, this sounds great to me.
The more interesting question is: can we train a Dota model that plays with all 124 heroes today?
AI's APM was limited to a level lower than pro human players. If they had allowed micro heroes AI would have a big disadvantage.
That was the most interesting part about the AI Dota game. AI isn't just better than humans at mechanical level. Even more surprisingly, AI (at least Open Five) isn't significantly better at last hitting than pro players.
Though to be fair, the human players had to rely on muscle memory to win lanes (CSing, blocking waves, pulling, trading hits, cutting waves, stacking, etc.); whereas the AI could perfect the timings down to the fraction of a millisecond.
In a similar vein, it would be fascinating if the AI had to also evade bot detection, that is appear (nearly) indistinguishable from a human player.
In the next game it played they had made it react even slower and then it no longer beat tournament teams.
See @deep blue for more. Or, any strategy game made in the last 20 years with AI difficulty mods.
For example, the more realistic human character animations get, the deeper we seem to fall into the uncanny valley. The fact is that humans themselves naturally move in ways that look weird when put on a stage, so mocap tends to look sillier the better it is. Which is why we have theater, where movements are exaggerated.
Anyway, with AI characters, I expect it will actually be more frustrating and boring than not if they have realistic lives and schedules. All the littler irritations that we deal with and accept form real people just become friction in a game. Games, as movies and books and shows and plays and illustrations, don't need to be more realistic to be better. Media is caricatures of real life, with important information intentionally presented to give us a good experience. Taking inspiration from real life can give us better mechanics but blindly mimicking real life will give us shitty games.
I think it depends on the type of game you have, but I wouldn't underestimate this type of technology for say, open world games where it might make the game more immersive due to convincing realism.
I think you misread my posts. We don't have awkward animations because our mocap isn't good enough, we have awkward animations because typical human motion looks awkward - our brains just mostly ignore that.
People are awkward; we don't actually want characters in games/movies/etc to be like real people. Very few movies, for example, would be well served by conversations frequently and for non-plot-related reasons being interrupted by loud noises, having people talk over each other and nonverbally try to figure out who gets to speak, having characters ask "What?" and then begin to reply without waiting for the answer because their brain caught up half a second later, etc.
But I don’t think I can agree with what you meant then - why would our brain mostly ignore it in real life but not in video games? Where is the transition from something feeling real (in an immersive way) to us not liking it because it feels awkward, and why does it happen? I’m imagining a “perfect simulation” game which is like real life in ways that matter/don’t get in your way in terms of gameplay - I think everyone would be awed (of course this can be argued though). What would need to degrade in terms of realism for it to seem awkward/not be immersive anymore?
I agree with the movie example, but in a game you don’t have to watch the mundane - it’s just background “noise” to make the world believable.
I remember Ultima V in the 80s had this. NPCs had their own daily routine and went here and there throughout the day. A couple of side quests relied on this mechanic—you had to learn their schedule and catch them somewhere at some time.
I'm pretty confident that people will find ways to make fun and interesting games using them, just like they have with every other computer technology over the last 40 years.
You'd be surprised at how untrue that is, unfortunately.
> I'm pretty confident that people will find ways to make fun and interesting games using them
I agree - but also, that's a different phenomenon than simply inserting AI naturalistic characters and expecting them to be fun to interact with.
Games bend over backwards in a million ways not to be realistic even when they seem realistic: from silently helping your odds, to intentionally breaking physics (see: coyote running in platformers), to helping your movement align with enemies.
Even the most realistic games with human locomotion almost all rely unrealistic ankle-breaking acceleration changes for your character because you'd feel like you were piloting a bowl of soup otherwise.
The reason we don't have full agentic NPCs right now isn't a technical limitation: it just wouldn't be fun.
Usually you're the hero going against impossible odds with NPCs, in the real world you'd just die.
Games like GTA constantly have NPCs saying outlandish but comical quips at a rate that far surpasses real life.
Hard coded NPC behaviors in games often become the defining features of those games
-
There's nothing wrong with realizing that outside of very narrow context like simulator physics, people generally don't seek out realism in our entertainment. People want a certain degree of escapism that relies on breaking the rules a little, or in the case of NPCs a lot.
The questions is not, "Is there a level of realism that is unfun," because duh, of course there is. The question is, "Will new advances in AI allow game developers to bring life to NPCs in a way previously not possible, that's also fun." And duh, of course it will.
Pointing out ways that increased realism can be unfun does nothing to prove that there are no ways in which it can be fun. Y'all are betting against the creativity of the entire game development industry. Absolutely wild.
I expect AI will be used to scale dialogue options, NPC model assortments, NPC voices, etc. to an amazing degree too at build time.
But realtime "Large Transformer Models", to generate how NPCs act is almost a "duh no one will want that", when you can get 99% of the benefits by baking the output of gen AI in the game with none of the drawbacks like losing the shared experience of a game or playing roulette with the quality of someone's experience.
AI would let you bake in insanely intricate dialogue trees today
You're making tremendous assumptions here:
- That the shared experience of game is always more fun than individual experiences.
- That only 1% of benefits can come from real-time AI decision-making.
- That using AI will mean playing roulette with experience, AND that playing roulette can never be fun.
I would bet all the money on earth that game developers will prove you wrong in each of these areas over the course of the next 3-5 years.
At first, devs start using generative AI for NPC interactions. It forces you to keep Internet connection and it feels extremely laughable. Everyone hates it.
Then 3 years later it becomes so natural that every AAA game has generative AI, and no one even mentions that as if that was always how games were made.
Can you name any games built around the Y-combinator, CRDTs, or PCA? (Please oh please let there be games about principal component analysis.)
That being said, I've noticed that realism in games has a function beyond gameplay: it makes the game cooler.
For example in Rainworld, the enemies keep doing their own thing even when you aren't around. This has some gameplay value (makes the game less predictable), but I think the main value is that people talk about it and think it's interesting. So it changes people's perception of the game world in a positive way.
Another game that has this property (and that has somehow escaped the befuddled reviewers' ire) is Don't Starve. e.g. I once got chased by the Deerclops and ran through a Beefalo herd. After a while I realised the Deerclops wasn't hunting me anymore. Next day I went back to the Beefalo herd and there was a huge amount of meat and Beefalo wool, and a Deerclops eye, lying around.
After playing games like that, I play e.g. Hollow Knight, or Zero Dawn, where every enemy you kill comes back the next time you revisit the area and it feels so fake.
- Ability to reload when you fail
- You can choose your gender
- Actually, be anything. A dwarf, elf, dog, tall, short, muscled, etc...
- Novel physics (magic)
- Sooooo much more
"context": {"Go about your daily life, but you have an IQ of 50 and have zero ambition."}
Crafting instructions that lead to NPC's working to achieve a goal that wasn't anticipated, and provides you with a resource that couldn't otherwise be had in the game, will be a whole new aspect to gaming, and inspire a desire to experiment and explore.
Seriously, have you no imagination? Why sit around coming up with reasons it won't work instead of ways to make it fun?
If you're making a highly interactive, dynamic game, you don't even need detailed language for NPC interaction. You may as well use simple templates or even symbols.
> Then we'll make new kinds of games where the unpredictability of the NPCs is a core mechanic.
This is impossible, you need the NPC to be predictable on some level to make a fun game. Even unpredictable NPC needs to have a predictable personality on some level, total randomness isn't fun. Like, imagine a terrain generator that just randomizes terrain on each tile, that wont be fun at all, that is what a basic random personality would be like.
Think of a human opponent, they are very predictable, just looking at a human player and what he does and I can predict what he will do in the future. Not perfectly, but players aren't that random. To make an AI that feels good it has to be very predictable.
The main problem with "smart" bots is that they have so far always been way less predictable than humans, they get a strange edge cases and bugs where they start to act very dumb and strangely, that feels like a bug to the player and isn't fun. Or their smartness makes them do the same thing every time making them even more predictable than basic scripting, either way they are worse than basic scripting.
Getting over these issues is a really hard problem, LLMs hasn't helped solve that so far.
tongue in cheek counterpoint: Rimworld players love Random Randy :P
I think it really depends on the game though, but you're right 100% random in an RPG could be really annoying.
Right now I'm into games like Factorio and Captain of Industry and they've both recently had blog posts about how they do terrain generation and CoI stuck out because you can manually plop features like mountains and then it procedurally generates the mountain range[1].
There's been a lot of games recently that seem to be doing procedural land generation, is there not a way this can be applied to AI personalities as well or is there no overlap between them? It kind of feels like procedurally generated personalities should be do-able but it sounds like there's something more going on that complicates that?
Even randy random isn't entirely random, people love it since it sends you big threats, so it is coded to ensure it throws you big threats. If it randomly didn't send big waves people wouldn't like it as much.
"If Randy has not fired a major threat after 13 days, the next Randy fired event becomes a major threat."
https://rimworldwiki.com/wiki/Randy_Random
> There's been a lot of games recently that seem to be doing procedural land generation, is there not a way this can be applied to AI personalities as well or is there no overlap between them?
I'm certain that is possible, but we don't have nearly as much intuitive understanding how to generate full fledged personalities hooked into an LLM that changes how the character acts and his motivations etc that will actually work well when put in a world and interacting with other NPC's in that world.
Terrain is just really easy to generate well enough, almost everything else is way harder.
Diplomacy in games have always been terrible because it just comes down to preprogrammed numbers. Humans might be the same at the very base level but there’s no way we could code that all out.
Now you could truly make those leaders act like they “should.”
Surely you've met some of these pessimists before.
On the other hand, I've loved co-op games. Maybe (it's a big maybe) I'd enjoy an AI driven NPC co-op character. It assumes they can get them good enough to be fun to play with, do what you want, etc. Player: "Let's go to the castle", NPC: "See you layer I'm going to the bar". You either follow the NPC or play without it.
But again, that's where LLMs can thrive! Currently if I want to play anything co-operatively I have to either convince my partner or friends to pick up the game and care about it, or hope that the existing community of players is such that I'm not spending what can be a very uncomfortable time trying to relate to kids/teens or toxic adults. If all the NPCs were driven by LLMs whereby I could modulate their play-style, personality, drive/ambition I could be looking at genuinely stimulating, creative, and original "co-op" experiences from bots.
The cat-and-mouse game of stopping these goldfarmers just became exponentially harder.
During the teenage years while the neocortex is growing, teenagers are practicing and honing fine motor control. Video games help develop that, as well as learning about social interactions, emergent system behaviour and strategic vs tactical thinking.
I’m not saying they don’t have downsides, but there are some upsides too.
Sure, if you’re an addictive personality using video game escapism to ignore your life problems, that’s a whole different thing (and even without video games this type of person would just find another form of escapism)
So I don’t agree with your generalisation of games as “time wasters” - maybe for you they are, but not for everyone, I don’t play them anymore very much (just a bit of chess every now and then) but they provided me with a lot of understanding in my formative years
A large percentage of social interaction in a game like world of warcraft is profoundly negative and maladaptive. I'm not sure I would want my child learning about that in a MMO.
We don't give kids as many opportunities to make mistakes in real life, I dunno.
I also agree that even MMORPGs have their upsides. But as a genre they're pretty unhealthy.
We've reinvented how arcade games used to try and extract maximum quarters, but the iteration cycle is so much faster now that we can't really play whack-a-mole on all the pathological human manipulation strategies people deploy now, and with people not being able to physically walk away from their phones or other devices in many cases, it's Bad(tm).
I used to think: “Anything that got rid of the ‘pointless grinding’ aspect of RPGs would be an improvement.” And then micro transactions and pay-to-win got invented. I didn’t think it was possible but game designers somehow actually managed to make RPGs even worse!
Some people find incremental progression for weeks or months or years at an achievement point or something similarly ephemeral to be satisfying. Some people think that it's not worth doing anything that you can't get in one play session.
Neither of them are wrong, for them, necessarily, but I do believe there's a nonzero amount of friction you want to introduce, and various sweet spots for friction versus people saying it's too much but still loving the game and playing it versus it actually driving away too many people.
I have played a number of fairly grindy games to various definitions of completion, and found it satisfying to play them more than once and get a lot of the optional objectives again. I have other friends who spend a week just grinding out the cap of various items in early game and then coast through the rest and think anyone who thinks a week of nothing but grinding to do that is mind-numbing and unpleasant is just wrong. I have other friends who have refused to play games older than PS4 era because the graphics not being hyper-realistic drives them out of the game and they think anyone who enjoys anything remotely grindy is just rationalizing broken systems.
But ultimately, for many games, you play them for the incremental satisfaction, and ensuring you get a steady drip of that and incremental fixes of larger satisfaction until completion, without turning each hit into a microtransaction exploitation, is the essence of making a game that people want to play again. Nintendo does this very well, in a number of their games even in modern times.
I feel like not understanding that there is a certain amount of padding required for you to enjoy a game as much as you might otherwise, because humans don't process things instantly or handle a truly constant flow of engaging things well, is one of the things a lot of people I see who play games and complain about being bored when they bounced off spending more than 15 minutes on it don't appreciate.
(Not saying you're one of those people or making that argument, just that I see a lot of people arguing that various kinds of friction could be removed, without thinking about how much that changes the perceived experience.)
I didn't say they were
Not sure how it'd compare against similar amounts of youtube/netflix though.
How? A lot of games could be seen as a sort of 3D chess.
I have fond memories of LAN parties growing up, where socializing was as big a part as the actual gaming - it's not like we were sitting there harvesting wood for hours on end!
Much more than chess which is mostly a individually played game whereas an MMO is a cooperative game.
- Paul Morphy
One of my favorite chess quotes. As an avid chess player, I can't agree more.
It feels fun to be rewarded for something you accomplish in-game.
In many singleplayer games, you can slide difficulty up or down to change the effort:reward ratio.
In an MMORPG, though, you have different groups of players with different amounts of time. You want to make it fun for both the kid on summer vacation who is happy to spend 80 hours a week on a game (not a choice I want for my kid, but I was a kid once too) and an adult who has a 60-hour work week and exchanges 2 hours of sleep after the kids go to bed to play.
That means the person with more money than time will want to buy things from someone with more time than money. But this causes all kinds of distortion in the game balance and economy.
I don't know that this is solvable, whether you're trying to balance against cheap labor or AI bots.
Now playing alone for the dopamine rush of successfully grinding repetitive tasks: yeah, that's a bit of a time waster. Maybe therapeutic for some, and definitely not the most harmful way to spend time and get validation, but also a bit pointless. But I would argue that if you play an MMORPG alone you're doing it wrong. If you don't have friends at least get engaged in a guild and spend countless hours improving real-life social and leadership skills.
But point is, that realization isn't that simple, it took a long time for cooperative games to become common. In early days game consoles had cooperative split screen to let two players play at the same time, not because that was more fun, so it took a really long time for cooperative modes to become standard in online gaming because it wasn't at all obvious that people liked cooperative play.
MMORPGs were the main cooperative online games for a long time. Today we have dedicated short session cooperative games, those are still very popular.
Speaking of social cues, interacting with others specially in a complex environment where there can be severe competition as well as cooperation and difficult coordination, is something that also is worth practicing.
I have nothing against solo games, but this kind of thing is not practiced in a solo game.
Finally, I think other kinds of games (e.g. in competitive games) tend to have very simple interactions and objectives, compared to an MMO: there's a clear objective to win that's shared by everyone. Some MMOs have much more interesting interactions, where each person is interested in a different thing, and I think this contributes to a very rich atmosphere that isn't just 'Go win, try to win match, go out', i.e. more life-analogue (without other limitations of life, like you can't actually die, and being poor isn't as terrible as it often is IRL :( ).
The inputs to a computer game are more limited, you can't see people, their faces (and sometimes voices), the graphics are still a far cry from the more beautiful places.
Also, real life is full of responsibilities and large parts of it still, well, suck (bad jobs, exploitative practices, etc.). I think we're improving somewhat (greatly hampered by greed and power games).
If you have interesting activities IRL, like a great fulfilling job and hobbies (that are also potentially useful in other ways, like charity work), then by all means, but I think virtual worlds have their place in our lives.
"Why did you do that if it doesn't make money?"
So that's what was going on in the Matrix when the humans were staring at all that green text.
NEO: Do you always look at it encoded?
CYPHER: Have to. The image translators sort of work for the construct programs but there's way too much information to decode the Matrix.
Everyone learns to read the raw code because humans don't have access to enough compute to decode the live matrix data stream
Like, Runescape was already distilled into a surprisingly good idle game in Melvor Idle. You could take a slightly different path where the "idling" is instead a matter of programming and resource allocation.
Then you may be interested in Screeps: World
act(screep) for screep in all_screeps // Independent evaluation
act(all_screeps) // Global coordinatorI think what people are missing, is that by the time you build a controller interface, and a screen grabber, and have an AI that can interpret the screen grab, understand and play the game, that this is super incredible, and really the humans are probably already being herded into Soylent Green processing centers to feed the remaining humans that are kept around for maintenance tasks.
And this post is showing you an AI that can look at the screen grab and play the game.
Here to offer praise to the AI overloads. Hope they read my comments later and know I was a true believer and should be included in the maintenance crews they allow to live.
And if that ever becomes a real problem they add a 3D camera to the controller or require it to be positioned below the screen, and use ML to decide if there's a human in front of the machine or not. AI cuts both ways. Integrated hardware devices aren't necessarily easy to just cut into. That's why it's hard to mod consoles to begin with.
If that's true, it must be easily bypassed because keyboard/mouse-to-controller adapters, like the XIM MATRIX, are still around after over a decade and work with the latest consoles (e.g., PS5).
Google ruining every part of the online experience, piece by piece
“ SIMA agents trained on a set of nine 3D games from our portfolio significantly outperformed all specialized agents trained solely on each individual one. What’s more, an agent trained in all but one game performed nearly as well on that unseen game as an agent trained specifically on it, on average”
Many people I talk to assume that when a LLM gets something right, it’s because that specific thing was in the training set. Although the experience of human transfer learning is intuitive to people, I find people have a hard time appreciating that it can happen in algorithms too.
That's typical of the poor generalisation displayed by neural nets and clearly not how humans do transfer learning.
Then I got the creepy feeling that this is likely the kind of wargames that are already being tested. We'll probably also need reverse safeguards where the AI raises concerns and requires confirmation before carrying out some requests.
So if I could get reports like "+10% failed to understand how to chop their first tree" that would be good.
Maybe we simply keyword matched on "video games" and "simulations". Or, perhaps more cynically, we're foreseeing a future in which AI agents don't care to differentiate between shooting at the enemy combatant in Call of Duty verses shooting at us in real life.
Some more discussion on the official post: https://news.ycombinator.com/item?id=39691783
If you think about it abstractly, humans are basically models that take input from our senses, do some internal processing of that and then take actions with our bodies. SIMA is the same - it takes input from video, and takes action through keyboard actions. There is nothing against introducing additional types of input and taking different actions.
The ability to train on one game and transfer that knowledge to a different game should allow future models like this to train in games, by reading text, watching videos etc, and then transfer all of that knowledge to the real world.
Here we go ladies and gentlemen. There's only one direction from here and it is that these things will keep improving and becoming better...
The goals of those niche bots were certainly different but in some ways the recent hype doesn't surprise me as much having experienced that period too.
I reflect that the limited internet then drove innovation in offline play that had really stagnanted till recently, I'm looking forward to the first game that really pushes the limit with their NPCs using some of this new tech
Video games might be even harder to predict sometimes, since the physics simulations have very strange edge cases. There are no physics glitches in the real world.
Massively, in a game running into a wall is perfectly normal and valid strategy to get close to it, in reality that will wreck you.
Or running on a fist sized rock has no consequences, in reality that destroys your foot. Reality is full of such extreme threats everywhere even in normal homes.
Video game physics can also be made more difficult - it's not a stretch to think about a reality simulator that dials these real-world effects up to 11 for AI to train in, at which point you need the meatspace bot and many things start to happen.
In terms of alternative strategies, Google DeepMind also has an amazing robotics team with lots of fantastic work for real-world robotics - including multi-robot generalists, showing positive effects when co-training one agent or model on multiple environments/bodies. Their prior work was very inspirational to us in SIMA! https://deepmind.google/discover/blog/scaling-up-learning-ac...
What would happen if, for example, the system were trained on simple 3D puzzle games, with natural language instructions such as “solve this puzzle” with a hint like “X or Y strategy might work”? I see that menu navigation is part of the training set. Can this thing learn to “read” or does it just learn the results of menu actions? If it can learn to play No Man’s Sky.. can it learn to play Zelda?
Tesla has an alternative. If you can get your devices widespread and can be recording observations and actions, you can collect huge datasets in the real world.
The first embodied, vaguely general, multi-task, useful and economical robots might just really open up the virtuous cycle of experience, learning, feedback and improvement. If I had to guess where it would come from right now, I'd pick Amazon warehouses.
this is direct violation of google ai principles on autonomous weapon development: [1]
[0] Screenshot from SIMA Technical Report: https://ibb.co/qM7KBTK
I dont think an agent fighting in a video game really counts? There is quite a significant gap between an FPS and a missile launcher, and it would be a waste not to explore how these agents learn in FPS environments.
They intentionally included combat training in the dataset. It is in their Technical Report.
How can combat training not be interpreted as "principal purpose or implementation is to cause or directly facilitate injury to people"?
Do you believe the agent was trained to distinguish game from reality, and refuse to operate when not in game environment? No safety mechanisms were mentioned in the technical report.
This agent could be deployed on a weaponized quad-copter, or on Figure 01 [0] / Tesla Optimus [1] / Boston Dynamic Atlas.
[0] https://twitter.com/Figure_robot/status/1767913661253984474?... [1] https://www.youtube.com/watch?v=cpraXaw7dyc
From the very beginning DeepMind has promised that its advances in getting agents to operate in virtual environments (mainly, games, e.g. Atari, or chess-go-shoggi) will lead to agents capable of operating in real-world environment with comparable skill and capabilities. So far, it hasn't happened and it keeps not happening. Now they added some super-trendy natural language stuff in, and it will keep not happening. If the point is to get computers and robots to not be dumb in the real world, that's all just a big waste of time and money.
I feel like in terms of game dev nothing supersedes state machines.
I know there's Gymnasium: https://gymnasium.farama.org/ and CleanRL and plan to mess with those but I havent seen much evidence they can be used to make a compelling adversary AI.
Gaming physics change all the time. I find it hard this "generalist AI agent" can play with a constant shifting gravity in XYZ for example. It does however offer a good training space for real world where gravity and other physics mechanics are consistent.
Can we hook this thing to a reaper drone and have it follow natural language instructions?
I can't find a good reason for computers playing videogames. I read another comment saying that they could be your buddies in an adventure game... what's the point? The fun is to play with other people. We already are able to play with bots (different algorithms rule them), so I can't see why someone would prefer this over them.
About traslating this from a virtual world to the real world... I can't imagine who would think it's a good idea to give this type of freedom to machines in a physical world, were consequences are way riskier than something digital (and yes, digitally they could empty your bank account, physically they could kill someone. One is much worser than the other).
Right now there are robots in many factories around the world, some are discrete machines that aren't tethered down and have movement capabilities. You don't think that there are factory managers/etc out there drooling about the idea of getting those or something similar to be able to do general factory tasks?
Imagine your employee who tapes up boxes before shipment quits one day out of the blue. "Hey, package carrying bot 9000, can you go tape those boxes? I'll have someone show you what to do"
Not necessarily a good idea still but just because we don't want it doesn't mean there aren't a million beneficial uses of this kind of generalizing.
I think strategy might be more important in the future than switchy aiming
There may still be holdouts, but the society will largely see them as old curmudgeons resembling the people of today who claim that "humans are incapable of maintaining long distance relationships, and you're only fooling yourself if you think you're in love". Or as a more extreme example, people who maintain that interracial or intercultural marriages are inferior to marrying in-group, because "shared experience" is fundamental to love.
All speculation, of course. But I'm a firm believer that humans are foolish and flexible - we're very easily fooled by anthropomorphic things, and AIs are the ultimate anthromorph. Our monkey brains don't stand a chance.
Extreme solipsism and loneliness in reality, maybe not experienced but in reality.
No one will have time for the chaotic and very human traits that is the essence of actual human existence instead craving the super engineered dopamine machines AI will become.
They really should not ask any AI agent to shoot at anything. Especially when it's not very good at it.
https://deepmind.google/discover/blog/sima-generalist-ai-age...
This one got points slightly more quickly for whatever reason. Usually in these cases of duplicate submissions dang will merge the stories together.
To the DeepMind team I would say that a snappy summary accompanied by a video at the very top of the page instead of a bunch of whitespace and a static image you have to scroll past would likely help the blog post be more viral, as this tweet demonstrates.
With a made up title which you should avoid, it's in the thing https://news.ycombinator.com/newsguidelines.html
please use the original title, unless it is misleading or linkbait; don't editorialize. There are regular moderation comments about it, they just say the opposite of what you think they say
https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
https://deepmind.google/discover/blog/sima-generalist-ai-age...
not creepy at all.
And once validated, sell to the military?
> Ultimately, our research is building towards more general AI systems and agents that can understand and safely carry out a wide range of tasks in a way that is helpful to people online and in the real world.
This makes me nervous.
I hope AI agents that take actions in the real world are regulated at least as much as self-driving cars have been over the last decade. Or at least AI agents that interact in public spaces.
> Ultimately, [..]
I swear this short paragraph style rounding it off with an “ultimately”, “in conclusion” didn’t use to be so common. :
Ai is already strongly influencing how people write. After being successfully deployed for a year.
DeepMind> kill dissidentsCan't wait to see the leaked footage of war crimes showing robots murdering civilians and teabagging their corpses
If it can be done in 3D virtual environments and video games, it shouldn't be much of a leap to do it in the real world. After all we have cameras, voice recorders, sensors, etc that can map the real world into 3D virtual environments already. Have they tried linking this generalist AI to a robot to see how the robot does in the real world?
The project required lots of reverse engineering on my part to make a web-based facsimile of the game such that it's possible to conduct controlled experiments on the language capabilities of current agents.
Hopefully what I've created will be useful for others, because unlike big tech, I've released all my code under the AGPL [2].
[1]: https://pl.aiwright.dev [2]: https://git.sr.ht/~dojoteef/pl.aiwright
Their approach is one that works for simple directives: "Go to ship" or "Pick up iron ore" which lends itself well to sandbox-like games (which seems to be a major focus looking at Deepmind's tech report). Similar research has been done in Minecraft [1].
These instruction following agents are more an RL achievement than a language understanding achievement. On the other hand, Disco Elysium has over a million words of dialogue, and solving the quests requires an agent to understand and reason about language much more extensively. People have looked at text-based game agents, like Microsoft's TextWorld [2], but these are much smaller in scope and not easily adapted for humans-in-the-loop.
My work bridges that gap, focusing on the language aspect, rather than navigating a 3D world. Again, they are definitely different objectives, but as a sole researcher there's no way I can compete with Deepmind's budget and manpower anyway. Just look at the extensive author list in the tech report. So it doesn't make sense to necessarily focus on outcompeting them in producing a better generalized RL agent (in fact I merely use GPT-4). Instead, I made a publicly available experimentation platform that allows others to be able to build upon this work, which is valuable for the community at large.
At least, that's my take.
[1]: https://sites.google.com/view/steve-1
[2]: https://www.microsoft.com/en-us/research/project/textworld/