Programmer Creates An AI To (Not Quite) Beat NES Games
techcrunch.com
techcrunch.com
""" On a scale from "the title starts with Toward" to "Donald Knuth has finally finished the 8th volume on the subject," this work is a 3. """
The Abstract is 16 words. The introduction begins with "The Nintendo Entertainment System is probably the best video game console, citation not needed." Need we say more?
Also, I wish non-programming fields of media would also adopt pythonic triple-quotation marks.
What overlying function in print would they serve exactly?
This really makes me want to go back to my game and play some more with AI. Although, at the same time I realize that's the mistake I made - making the game/platform for creating AI. It really takes a toll. This guy uses existing NES games, which allows him to focus on the AI and even generic game learning. Brilliant.
You can clearly see the advantages of AI in the short-term play (think in the order of milliseconds, exact frame-by-frame button presses) over anything a human could ever achieve.
Imagine combining the short-term button-mashing of said AI with the long-term planning of human players...
Feed your bot AI with information that is 200 ms old and force it to extrapolate 200 ms in the future for its aiming/movement/etc. logic. It will make the bot much more human-like with noticeable reaction time, so it could be tricked by sudden changes in your behaviour. Such a simple but powerful AI trick.
The video is long but worth all 16 minutes. I like how he isolated such a dead simple heuristic for success in video games. It's not surprising, in retrospect, that "numbers going up" would correlate with winning, but to prove it in practice is pretty neat.
If that makes sense.
For example, unpausing Tetris. Or having Mario run left.
Teaching a computer various concepts of 'crazy' must be tons of fun! :-D
The author is humorous, poetic, and brilliant.
War Games: Very, very basically, from deep, deep memory...
A computer about to launch a global nuclear war is played at tic-tac-toe to learn it cant win. It deduces that the only way to win a nuclear war is not to start the thing in the first place. Any other strategy is a lose, or what we call mutually assured destruction.
Its obviously more complicated than that. Even though its 30 odd years old, it worth a watch.
want to play?
But no.
"Murphy ran a few other games through it, including Tetris, and found that the program would eventually just pause itself rather than continue playing and lose, a tactic shared by annoying, over-competitive cousins around the world since 1985."
Direct link to video: http://www.youtube.com/watch?feature=player_embedded&v=x...
I do agree with you though.
This is why he mentions that Karate Kid didn't work well, because one of the factors of a successful play-through was that your opponents health is decreasing, which is something that this AI doesn't look for.
The ability to rewind time wouldn't make for a very interesting play-through of Battleship, but depending on how game state is stored, it might not even be able to accurately deduce a good game state from a bad one.
I don't believe that was the case. From the video he mentioned it using the power kicks up all on the first enemy because they produced the most favorable outcome for that game, at the expense of later games when more powerful moves are needed (since the AI can't plan that far ahead).
I can't watch the video at the moment but I imagine the paper goes much farther in-depth with the internals of this AI.
you should compare it with a game whose rules you don't know, being able only to understand your score. like trying to play spider solitaire on windows without knowing how the game works (it allows you to undo your moves).
notice that in your example of solitaire, knowing the values of hidden cards really changes the nature of the game and makes playing the game much less strategic
It is clearly an April Fool's hack.
This is from the paper. "I tried again, and it was a huge breakthrough: Mario jumped up to get all the coins, and then almost immediately jumped back down to continue the level! On his first try he beat 1-2 and then immediately jumped in the pit at the beginning of 1-3 just as I was starting to feel suuuuuper smart (Figure 7). On a scale of OMFG to WTFLOL I was like whaaaaaaat?"
The issue for modern games is the amount of memory used.
spoiler alert | You are travelling thru space encountering an ever-increasing number of what appear to be robotic ships with whom you can briefly interact before they end communication and attack you. Long story short: a planet-bound race, the Slylandro, purchased a robotic ship for exploration and to establish friendly contact with other races. The ship had the power to self-replicate; it also carried a small defensive weapon and a powerful lightning-bolt type tool for breaking-down and gathering resources for ship replication. The Slylandro wanted more probes faster so they turned the replication priority up to 11, overriding all other priorities. When a ship meets you, it attempts to initiate contact but is quickly forced into resource-gathering/replication mode & proceeds with breaking apart your ship into raw components. It is, essentially, a paper-clip maximizer. :)
http://radgeek.com/gt/2006/04/28/antieconometrix_comix/
It's too bad that game that need more lookahead - Tetris, etc, are not amenable to this approach.
It would be interesting to see how it responded when the 'Pause' function was disabled. It seems like it would need to track 'past future' state to find a solution.
Edit: Or maybe try to filter out 'constant' sources of increasing value. Treat the stacking points as noise.
Is the baseball player's anticipation of the future really all that different from what this program is doing?
Is a decade or two of baseball practice really all that different from a computer playing and re-playing the game over and over looking for optimal strategy?
I don't really see how it's time travel but that's what he called it.
Genetic algorithms (and variants) can pick up some very interesting patterns for achieving the target fitness. I don't think this is necessarily anything new, and GAs have their limitations, but applying them to multiple NES games is impressive. I'm convinced that GAs will be an important part of future computing.
Similar, but more interesting
But it's not even close to beating NES Games. I'm not a expert in AI but I'm not sure it actually achieves anything.
I'd consider it more an art project than a computer science project.
Look at that AI play. It's godly.
Also what is that AI doing when it decides to jump into the pit at 0:45? It clearly sees the path to jump over it, but it chooses the one that jumps into it. Watching it frame by frame it seems to figure out a path to jump out of the pit several times, but chooses not to, until eventually it finally does.
Either way though, I think the results would be sub-par.
It seems to succeed with move-to-the-right (position counter) but fall into what looks like Brownian motion on Pac-Man. It seems to me that putting momentum into its strategy (and changing when the objective function declined) might reduce the implicit branching tree.
What it simulates, then, is the tendency for people to fall into "a groove" and depart from it when their objective function ceases increasing.
But maybe the program failed to discover the bytes which hold the score, and so it didn't think eating dots was a particular notable activity.
Akin to making things out of Legos at best. (ie: "because!")
I can't find a reference but I recall an Emulator for NES that used to be able to be player 2 for something like 40 NES games.
This bot, on the other hand, understands nothing about the games themselves besides the input keys and the score number(s). It's playing blind.
4 directions and two buttons. You could point a crappy machine learning algo at it, and it would figure it out. That is far fewer variables to manipulate than what a Quake bot has to deal with.
I would bet money I could zap a pigeon every time it died in mario and it would learn to be a master in 12 hours. It won't master CoD anytime soon.
A Quake bot is easier to develop because it is written for a specific game.
I accept that bet for training a pigeon to play SMB :), although the scenario you described isn't comparable (it is only one game and you would be giving the pigeon extra feedback about how to play it. It still seems hard and I think it would be notable if you achieved it.)
There are Guitar Hero Bots taht "OCR" (your words) the screen to play the guitar. Not such a hard trick to read a score or see a turtle. (you know we have self driving cars right?)
The examples you gave are each specific to particular domains.
This whole topic reminds me of how it is much easier to develop natural language processing for use in a particular domain compared to generically.
Also "nintendo" is a particular domain. When you only have a "language" of 8 directions, A, and B, and only time to deal with in response to a screen that side scrolls in only one direction the programming is easy.
A chess engine has more potential decisions for a bishop than this has for a given move. And unlike the the NES, the chess engine has to adjust for changes in the behavior. Given the same input the NES would make the same choices. You can macro through most the games.
1. This was a paper submitted for a joke conference. It is in no way a giant leap in AI. It is something that some nerds find amusing. Like an anti-Rube Goldberg machine. It is trying to accomplish something quite complex using a ridiculously stupid approach.
2. Its stupid because it is looking at the memory state as just a list of numbers that change during the gameplay. It has no knowledge of the actual game. It doesn't parse the memory to figure out game state or anything. It doesn't even have any generic idea of what games involve. It figures out "oo, the number that starts at location 243 goes up in the training data. I'll mash buttons to try to make memory location 243 go up".
It's a cool project. If this is nothing compared to the things you've done, you're welcome to post them- we'd love to see them!
The emulator was presumably programmed to play those 40 games, too.
But the real measure of intelligence is to throw something at the AI that it's never seen before, and see how it performs. And that's the magic of this one: it's a general-purpose video game playing AI. You can find a new game that it wasn't specifically designed to play, and it will quite often do a decent job playing it.
The research here is not a computer program that plays games. Rather, it's a computer program that (attempts to) play any game, without having any understanding of what it means to play that game, purely by watching the buttons that a human player would press while playing that game.