>Humans play StarCraftthrough a screen that displays only part of the map along with a high-level view of the entire map, to e.g. avoid information overload. The agent interacts with the game through a similar camera-like interface
What exactly does that mean? Does it or does it not play by operating purely on image data human players would see on the screen?
How much of the system's interaction with game's interface is learned as opposed to hand-crafred and filtered through APIs?
It's amazing that most people here seem to think that system's ranking in a computer game are more important than its ability to learn from and interact with unstructured data.