> Atari games just from the raw pixels on the screen
It's important to distinguish between what sorts of games work well under this method and what sorts do not. Games that are variations of pole balancing, like Pong, fare better than more complex games like Asteroids, Frostbite or Montezuma's Revenge.
> Saying that AlphaGo is limited because it only knows to play one game, is like saying that humans are limited because Lee Sedol could only master at world level one game.
It's nothing of the sort. AlphaGo is a machine in the Turing sense. The neural network is a program that is the result of a search for a function specialized to playing Go. This machine, the program that the parameters across the edges in the graph represent, is logically unable to run any other program. Lee Sedol is a Universal Machine in the Turing sense, any statement contradicting this makes no mathematical sense.
> We limit software to specific domains only on account of efficiency, not because algorithms are fundamentally limited.
It is well known within the literature that these models do not make best available use of information when learning. They are exceedingly inefficient in their incorporation of new information. Issues include improper adjustment of learning rates, not using side information to constrain computation, having to experience many rewards before action distributions are adjusted in the case of reinforcement learning, samples per example in supervised learning. Note that animals are able to learn without explicit labels and clear 0/1 losses.
Humans and animals generally, even in the supervised regime, are vastly more flexible in the format the supervision can take.
For an example, look into the research on how children are able to generalize from ambiguous explanations as "that is a dog" and why difficulty in learning color from this kind of "supervision" shows just what priors are being leveraged to get that kind of learning power.
See here for an excellent overview of limitations in our current approaches to AI: https://arxiv.org/pdf/1604.00289v2.pdf
> "Learning without forgetting"
That's a great paper but it does this by minimizing prediction error drift by comparing before and post performance on the old task while learning the new. I do not know that this method will scale with increasing task numbers, considering Neural Networks are already difficult and energy-time consuming enough to train as is.