(I know the article says that it's been beaten, but it's wrong: it's improving over DQN by being able to explore 15 rooms rather than 2 rooms, but it hasn't cleared the first level, much less the whole game. An interesting breakthrough in how to define novelty in the ALE but not something as striking as AlphaGo.)
Credit assignment in a nutshell is "what actions helped me get reward"? For action games this is fairly easy - there are only a few moves between rewards. For puzzlers, something like left, up, right, up, left, left, left, left, up, up could get a reward. We can see there is a cycle in there which is probably not necessary, but maybe this was a much longer path than the ideal as well. Deciding which moves should get credit is a hard problem, but an important one. [2]
If you look at the results of the original DQN paper [3] you will see the games they fared best were ones where there were frequent rewards (e.g. Breakout). Things that are puzzle-like (such as Q-Bert) fared much worse versus human benchmarks, whereas action games like Breakout (which is fully observable given 4 frame context IIRC) were generally better than the human benchmark.
This paper seems to be a big step toward deep RL for more than just short term decisions and a huge jump towards goal oriented planning.
[0] Kulkarni et. al https://arxiv.org/abs/1604.06057
[1] Mohamed, Rezende https://arxiv.org/pdf/1509.08731.pdf
[2] http://www.scholarpedia.org/article/Reinforcement_learning#....
[3] Nature results are better but paywalled :/ NIPS paper here https://www.cs.toronto.edu/~vmnih/docs/dqn.pdf . http://www.nature.com/nature/journal/v518/n7540/abs/nature14... - Figure 3
On a somewhat related note, it seems clear that AI research and breakthroughs are occurring at breakneck speed. I wish there was a place where you could see expert commentary like your in layman terms on interesting or important papers that stand out.
Games like Montezuma's Revenge allow us to re-use huge amounts of already known information. We recognise there's a little person, and that's who we're controlling. We have expectations around what a jump might mean and that we probably shouldn't touch the skulls. We expect that moving off the screen to the next room then back goes to the first screen again. The game, in many ways, acts similarly to our normal reality (object persistence, motion, etc). We know we want to survive. We can even read the text on the screen and focus on increasing the numbers.
The AI has just has pixel values, and none of this information. It doesn't know about jumping, reality or skulls. It has a grayscale 42x42 view of something and is given a few ways of poking this world.
Edit - Perhaps another way of looking at it is this:
In go, you have 19x19 positions which can be in one of just 3 states. You must choose a move to make out of a large number of possibilities. This is repeated and the world changes very slowly.
In this game, you have 42x42 positions which can each be in a much larger number of possible states (somewhere between 8 and 128 I think). You have only a few possible moves but the world changes rapidly and sometimes completely. The interactions between the way the world changes in response to your actions are significantly more complicated.
For example, last year an AI beat NetHack (an incredibly complicated game that I think gamers can agree is "harder" than Montezuma's Revenge, although it's turn-based rather than real-time) for the first time. But the NetHack-winning AI didn't learn to play NetHack like a human would, and the prospect of that is incredibly remote.
Instead, the NetHack AI was hard-coded full of extremely detailed domain knowledge about NetHack items, maps, commands, monsters, goals, etc., and used search strategies to explore the dungeon and perform specified tasks given that knowledge. So it was barely doing any learning at all (although the dungeon map, starting inventory, and item descriptions are randomized on every play, so it did have to learn those things each time, and had explicit strategies for doing so).
The from-scratch success at video games is what most impressed people about DeepMind's original work; they were able to beat a whole lot of Atari games without telling the AI how to play. But if you see the level of complexity of those Atari games, writing an AI to play them would not have been such an impressive feat in itself. Although Montezuma's Revenge is a lot more complicated than something like Breakout, I think exactly the same consideration applies here. A computer could easily be the best Montezuma's Revenge player in the world already, but getting it there without encoding knowledge of the game calls for substantive new research.
AlphaGo is the program that beat Go. These game playing algorithms are variations of Deep Q Learning Reinforcement Learning algos.
Tech journalism isn't at its best when it can't distinguish between WOPR and a team of researchers.
So, DeepMind's greatest asset is they employ some really excellent people and have a substantial head start in terms of actually implementing AIs and getting them to work.