I implemented the DQN algorithm (used in this work) in Javascript a while ago as well (http://cs.stanford.edu/people/karpathy/convnetjs/demo/rldemo...) if people are interested in poking around, but my version does not implement all the bells and whistles.
The results in this work are impressive, but also too easy to antropomorphise. If you know what's going on under the hood you can start to easily list off why this is unlike anything humans/animals do. Some of the limitations include:
- Most curcially, the exploration used is random. You button mash random things and hope to receive a reward at some point or you're completely lost. If anything at any point requires a precise sequence of actions to get a reward, exponentially more training time is necessary.
- Experience replay that performs the model updates is performed uniformly at random, instead of some kind of importance sampling. This one is easier to fix.
- A discrete set of actions is assumed. Any real-valued output (e.g. torque on a join) is a non-obvious problem in the current model.
- There is no transfer learning between games. The algorithm always starts from scratch. This is very much unlike what humans do in their own problem solving.
- The agent's policy is reactive. It's as if you always forgot what you did 1 second ago. You keep repeatedly "waking up" to the world and get 1 second to decide what to do.
- Q Learning is model-free, meaning that the agent builds no internal model of the world/reward dynamics. Unlike us, it doesn't know what will happen to the world if it perfoms some action. This also means that it does not have any capacity to plan anything.
Of these, the biggest and most insurmountable problem is the first one: Random exploration of actions. As humans we have complex intuitions and an internal model of the dynamics of the world. This allows us to plan out actions that are very likely to yield a reward, without flailing our arms around greedily, hoping to get rewards at random at some point.
Games like Starcraft will significantly challenge an algorithm like this. You could expect that the model would develop super-human micro, but have difficulties with the overall strategy. For example, performing an air drop to enemy base would be impossible with the current model: You'd have to plan it out over many actions: "load the marines into the ship, fly the ship in stealth around the map, drop it at the precise location of enemy base".
Hence, DQN is best at games that provide immediate rewards, and where you can afford to "live in the moment" without much planning. Shooting things in space invaders is a good example. Despite all these shortcoming, these are exciting results!