Deep Mind Playing Montezuma's Revenge with Intrinsic Motivation [video]
youtube.com
youtube.com
That said, game theory has always been an excellent way to analyze AI systems. And using "modern" games (which generally provide attractive skins over a classic mechanic) certainly makes it easier to watch/sit through. When DeepMind starts beating people playing Diplomacy then we'll know we're in a whole new game.
AlphaGo is the program that beat Go. These game playing algorithms are variations of Deep Q Learning Reinforcement Learning algos.
Tech journalism isn't at its best when it can't distinguish between WOPR and a team of researchers.
So, DeepMind's greatest asset is they employ some really excellent people and have a substantial head start in terms of actually implementing AIs and getting them to work.
(I know the article says that it's been beaten, but it's wrong: it's improving over DQN by being able to explore 15 rooms rather than 2 rooms, but it hasn't cleared the first level, much less the whole game. An interesting breakthrough in how to define novelty in the ALE but not something as striking as AlphaGo.)
Credit assignment in a nutshell is "what actions helped me get reward"? For action games this is fairly easy - there are only a few moves between rewards. For puzzlers, something like left, up, right, up, left, left, left, left, up, up could get a reward. We can see there is a cycle in there which is probably not necessary, but maybe this was a much longer path than the ideal as well. Deciding which moves should get credit is a hard problem, but an important one. [2]
If you look at the results of the original DQN paper [3] you will see the games they fared best were ones where there were frequent rewards (e.g. Breakout). Things that are puzzle-like (such as Q-Bert) fared much worse versus human benchmarks, whereas action games like Breakout (which is fully observable given 4 frame context IIRC) were generally better than the human benchmark.
This paper seems to be a big step toward deep RL for more than just short term decisions and a huge jump towards goal oriented planning.
[0] Kulkarni et. al https://arxiv.org/abs/1604.06057
[1] Mohamed, Rezende https://arxiv.org/pdf/1509.08731.pdf
[2] http://www.scholarpedia.org/article/Reinforcement_learning#....
[3] Nature results are better but paywalled :/ NIPS paper here https://www.cs.toronto.edu/~vmnih/docs/dqn.pdf . http://www.nature.com/nature/journal/v518/n7540/abs/nature14... - Figure 3
On a somewhat related note, it seems clear that AI research and breakthroughs are occurring at breakneck speed. I wish there was a place where you could see expert commentary like your in layman terms on interesting or important papers that stand out.
Games like Montezuma's Revenge allow us to re-use huge amounts of already known information. We recognise there's a little person, and that's who we're controlling. We have expectations around what a jump might mean and that we probably shouldn't touch the skulls. We expect that moving off the screen to the next room then back goes to the first screen again. The game, in many ways, acts similarly to our normal reality (object persistence, motion, etc). We know we want to survive. We can even read the text on the screen and focus on increasing the numbers.
The AI has just has pixel values, and none of this information. It doesn't know about jumping, reality or skulls. It has a grayscale 42x42 view of something and is given a few ways of poking this world.
Edit - Perhaps another way of looking at it is this:
In go, you have 19x19 positions which can be in one of just 3 states. You must choose a move to make out of a large number of possibilities. This is repeated and the world changes very slowly.
In this game, you have 42x42 positions which can each be in a much larger number of possible states (somewhere between 8 and 128 I think). You have only a few possible moves but the world changes rapidly and sometimes completely. The interactions between the way the world changes in response to your actions are significantly more complicated.
For example, last year an AI beat NetHack (an incredibly complicated game that I think gamers can agree is "harder" than Montezuma's Revenge, although it's turn-based rather than real-time) for the first time. But the NetHack-winning AI didn't learn to play NetHack like a human would, and the prospect of that is incredibly remote.
Instead, the NetHack AI was hard-coded full of extremely detailed domain knowledge about NetHack items, maps, commands, monsters, goals, etc., and used search strategies to explore the dungeon and perform specified tasks given that knowledge. So it was barely doing any learning at all (although the dungeon map, starting inventory, and item descriptions are randomized on every play, so it did have to learn those things each time, and had explicit strategies for doing so).
The from-scratch success at video games is what most impressed people about DeepMind's original work; they were able to beat a whole lot of Atari games without telling the AI how to play. But if you see the level of complexity of those Atari games, writing an AI to play them would not have been such an impressive feat in itself. Although Montezuma's Revenge is a lot more complicated than something like Breakout, I think exactly the same consideration applies here. A computer could easily be the best Montezuma's Revenge player in the world already, but getting it there without encoding knowledge of the game calls for substantive new research.
And this still doesn't address my original question. A human who was able to master Montezuma's Revenge would have a dramatic advantage in learning and mastering the game Zelda compared to somebody who had played neither game before. What experience, if any, could this machine be expected to bring to Zelda, assuming no modification by the researchers?
> The ability to act in multiple environments and transfer previous knowledge to new situations can be considered a critical aspect of any intelligent agent. Towards this goal, we define a novel method of multitask and transfer learning that enables an autonomous agent to learn how to behave in multiple tasks simultaneously, and then generalize its knowledge to new domains. This method, termed "Actor-Mimic", exploits the use of deep reinforcement learning and model compression techniques to train a single policy network that learns how to act in a set of distinct tasks by using the guidance of several expert teachers. We then show that the representations learnt by the deep policy network are capable of generalizing to new tasks with no prior expert guidance, speeding up learning in novel environments. Although our method can in general be applied to a wide range of problems, we use Atari games as a testing environment to demonstrate these methods.
A human would instantly notice many ...
But isn't that exactly what Schmidthuberian intrinsic motivation is about? In Schmidthuber's account, which inspired the paper under discussion, intrinsic motivation is measured by the improvements (= better compression) to a predictive world model made by the learning algorithm.[1] J. Schmidhuber. Formal theory of creativity, fun, and intrinsic motivation.
So you use existing tools to encourage the player to survive, and then you add additional motivation to encourage the player to learn what it needs to do [ie explore]. And exploration is a common theme in games like Zelda.
I actually think that this technique will generalize to learning lots of games [eg the same engine might play Zelda & Montezuma's Revenge], but I don't think techniques in this vein alone will allow it to learn how to play one game and then immediately understand how to play the other.
[but "transfer learning" is a separate, challenging problem in AI]
Moreover, we already have machines that are extremely proficient at solving very complex games when given enough context: expert systems. One very notable example is the bot which successfully completed the game NetHack [0]. Would DeepMind's novelty-based reward technique work for NetHack?
Because you're absolutely right that a human who is good at one platformer will take very little time to adjust to a new one.
I don't know much about NetHack, but seeing that it's a Roguelike, I think these new techniques should definitely be tried on it or something similar.
Personally, I think that someone should be trying a DQN with an RNN rather than CNN in it to see if that helps on the harder levels. Or better yet, combine it with some of the memory mechanisms and see if it can start doing some real long-term planning.
What did the Wired editor mean by "solve" and "four tries"? (Or, for that matter, "complete"?)
within a fraction of the training time, our agent explores a significant portion of the first level and obtains significantly higher scores than previously published agents
... it's probably just the usual hype.
http://i.imgur.com/FfcXAmi.png
I am on Windows 7 using Chrome 51.0.2704.79 m and the text issue occurs in incognito mode as well.
Mon·te·zu·ma's re·venge (noun, informal)
Diarrhea suffered by travelers, especially visitors to Mexico.
https://en.wikipedia.org/wiki/Montezuma's_Revenge_%28video_g...