Learning ‘Montezuma’s Revenge’ from a single demonstration
blog.openai.com
blog.openai.com
After a bit of debugging, this appears to be a very intentional feature in the game rather than a flaw. That key appears after a while if you're not in the room (and don't have one).
Based on this disassembly: http://www.bjars.com/source/Montezuma.asm
Here's the relevant code with some annotations added:
I'm not sure if this is a previously known feature in the game (a quick google search does not reveal much). It would be quite interesting if the RL agent was the first to find it!
PS: If you launch MAME with the "-debug" option and press CTRL+M you can see the whole memory (atari 2600 only has 128 bytes!!) while playing the game. If you keep an eye on the byte at 0xEA you will know when the key is about to pop up. Alternatively you can speed things along by changing it yourself to a value just below 0x3F.
It is darkly humorous to contrast Hollywood's or scifi's "killer AI robots" that methodically hunt you down to these real world demonstrations of emerging AI. Maybe the first "killer AI robots" would exhibit similarly bizarre behaviors while they methodically hunt down the unlikely hero. :-)
On the other hand, if they are trained to optimize for energy usage, staying still when movement isn't needed can be an advantage.
But you can have a mundane, stable aircraft that is fly-by-wire as well. Like the Airbus A320.
If you consider the action state space, removing the 'do-nothing' state could provide a learning benefit. Consider a set of models that happen set the do-nothing state weights to zero, but manage to achieve a similar action using quick left/right movement. Perhaps these models train slightly better, meaning that in the same number of iteration steps they get to a better score than the models that do consider the do-nothing state.
Checking the video, you do see the person waiting from time to time. Perhaps this is an artifact of its demonstration learning episodes?
I do not see their Dota2 bot do this jitter movement. (Interestingly, the official/vanilla Dota2 bot does have this jitter!) This is likely because there is a benefit in being economical in your movements in that game: turning takes time. I postulate that an OpenAI bot for League of Legends, where turning is instant and free, would exhibit the jitter movements ;)
edit: Alternatively, inspired by the 'fly-by-wire' sibling comment: maybe spamming the emulator with left/right actions does provide a slight benefit. It wouldn't be the first time an AI finds video game exploits[1].
[1] https://arstechnica.com/gaming/2013/04/this-ai-solves-super-...
And even with a computer playing a video game (not the real world), your joystick hand gets tired.
That sounds very much like ED-209 in RoboCop (a film whose satire and themes in general are far more on-point and prescient than one might expect from the premise).
If there is some modeling of acceleration, it will be slower to stand and wait right next to a barrier or trap and not start moving until it disappears. Instead, it's best to have some speed already and be positioned so that you're about to run into the barrier on the frame before it disappears. Then on the next frame, where it does disappear, you move into the vacated space without slowing down or taking damage as the case may be (or starting from a stop, as you would have if you stood still while waiting).
If you're not already familiar with tool-assisted speedruns, I highly recommend tasvideos.org
In some TASes, the player will move/shoot/jump/etc. in ways that seem bizarre but are actually done to favorably affect the pseudorandom number generator (e.g. prevent enemies from spawning or shooting, so the additional sprites don't cause lag frames). Since the goal of this bot was score, not speed, there's probably not much of that going on here. I really like the move at 0:35 in the top video for Montezuma's Revenge, where the bot climbs up to the next screen and then immediately back down, presumably to reset the position of the green bug thing and get on the left side of it. That reminded me of some ladder action I've seen in Mega Man TASes.
And then there’s movements like the arm pumping in Super Metroid, where it’s done “because you can” while performing long stretches of movement that aren’t very entertaining/complex on their own. (One might say that it’s done for a demonstration of skill when human players do it, but there’s no similar justification for a TAS doing it.)
Arm pumping actually moves Samus forward a pixel each time. It was considered "too hard" for human runners, until I think hotarubi's 0:32 in 2006.
Real humans that play starcraft often have bizzarely high Actions Per Minute (APM) scores in the 200-300 range, and often most of that movement is spamming the same commands over and over.
Consider how that clicking would plunge if there were a single "kite move" command for units.
I believe the command spamming /u/tomlock is referring to is the phenomenon where players issue as many or more commands when doing nothing as when an intense sequence is happening. This is primarily done to keep rhythm.
Rather, a fair portion of that churn is due to an inefficient communication system or control structure between the player and the game.
It's analogous to pitting an android against a human at racing an old car: Both of them will have a high volume of activity with their left foot on the clutch and right hand changing gears, but it doesn't reveal anything amazing about their thought-processes. It's simply what manual-transmission cars require in order to accomplish a certain task.
I'm not sure what this statement is based on.
So they solved this by feeding the AI with a human demonstration, but have there been any attempts at giving the AI an explicit reward for maximising the "novelty" of the input state (i.e. the image on the screen)?
The game does not give the player points for reaching new rooms, but if the AI was rewarded for producing the "novel" state of a new room, then that would give it a drive to explore. Similarly, there would be an implicit penalty to the AI for repeatedly falling off a ledge or returning back to a room it had already visited (although some amount of back-tracking would no doubt be useful), whereas reaching a new part of the screen (by climbing a ladder, say) would be rewarded.
There are times where the AI would have to be patient and wait, but the window could be learned or set as a hyper-parameter. This might be enough to stop the unproductive behaviour of it jittering left and right continuously, since doing so does not produce a new state, relative to just standing still at least.
I think it applies to many games: you typically progress through different screens/levels/whatnot on your way to completing the game.
Calculating similarity of images is quite a well-understood problem, but you're right that generalising the idea of similarity across all types of input data, in a way that is helpful for the AI to learn from, and efficient to calculate, may end up requiring a lot of coding that's specific to the individual use case.
[1] https://paleotronic.com/wp-content/uploads/2018/05/5.png
Interesting that their approach didn't work for Pitfall (never played Gravitar).
https://news.ycombinator.com/item?id=17460392
I assume the agent somehow found this out and developed the behavior of going in and out of the room until the key shows up (which, with enough agent randomness it apparently will).
Iow, it was showing what beating the game would look like at some level of granularity. I guess the next obvious question is how far up you could dial the granularity and result in the AI still learning how to beat the game.
So they took that single play-through, chopped it up by room, and trained each room in reverse.
Or in other words- use the Domain Knowledge, Luke. Quit trying to learn everything from scratch. Because that's just dumb.
How well would this adapt if the map/layout changed then?
People beating this game do not do it based on a let's play video.