You wouldn't want AlphaGo to have to input it's commands using robotic hands right? It's the same thing in that, sure it might be interesting, but that isn't what we care about. Image processing and robotics controls are largely solved. Showcasing that a model can gain the ability to plan and think is the novel stuff here, and is the path to where "artificial intelligence" if any appears. That's the ultimate goal in playing any of these games.
Image processing and robotic control are very far from being solved problems. I guess you are saying that in the case of alpha go it would not be a super difficult step to have a camera and robotic hand physically move pieces around, and that's probably true. But I think in the DOTA case are new image processing challenges that interact with the AI in interesting ways.
I'm mostly talking about the need to move the game's camera around to gain more information. If you don't see your ally on your screen and need to see how they are handling a gank or something (full disclosure I don't play DOTA at all this could be a silly scenario). Then the AI would have to recognize this and move the camera to the allies location in order to gain that information. So really the novelty here would be in the network to somehow realize what information it needs and then further to learn how to gather that information. I honestly think that sounds like an extremely difficult next step.
That said, everything they've done so far is absolutely incredible (especially now that the AI can draft!!)
Humans dont concentrate on the whole screen, attention is directed...
I would love to see camera/mechanical interface like mentioned by others. Similarly, like you said humans don't focus on the whole screen. I would love to see how well the AI could perform if it was given something like blinders where only a small portion of the screen is in focus at any one time much like how human eyes work.
Otherwise, you still have to add time to react plus mouse travel time.
Basically, they let an AI loose on a simplified version of Quake's Capture the Flag. The AI processes game video output only and has learned several key strategies. The latest update has the AI with a winrate of 71% against top humans. Unlike the DOTA match the AI has no restricted reaction time.
The AI seems to be jittering the camera left and right to reconstruct a 3D image reliable from the screen which is aquite interesting way to compensate for lack of 3D vision (and the compensation our brain is capable of naturally to get a 3D intuition from a 2D image)
If this was actually the goal, they would add further mechanical restrictions beyond the 200ms delay to simulate the way humans players play. That way, human and AI would be on a roughly even mechanical playing field, leaving the differentiating factor only strategy/tactics.
As it stands, it looks like their victory is as much based on raw mechanical superiority as it is strategic/tactical superiority. Computers being able to be pixel-perfect accurate at all times, issue commands at ludicrous speeds, etc. is kind of an uninteresting advantage in the context of building strategic AI.
So this current ai is uninteresting because bots can always instantaneously begin to react on any feedback, whereas humans have to pan and drag the camera around to look at different feedback in the first place, let alone react. Mechanically, humans also have to move the mouse all over the place and think of key combinations, in addition to reacting. Not just clicking a static box on cue.
It -would- be interesting if bots were limited just like humans to the camera view, -not- an API that continuously feeds them information. The bot would then have to learn how to prioritize working the camera, and it would be limited to only what the camera sees, etc.
Computers already have perfect memory and recall, so when the image recognition tech becomes good enough to only rely on the visual input, are you then going to say the bot must now limit its recall to "human" levels?
For many things in Dota you also need to move the mouse cursor to a specific point on the screen which obviously takes longer than just pressing a button.
In chess, both players always have perfect information about the game-state, and this is far from the case in Dota. OpenAI does account for fog of war, so it's not COMPLETELY omniscient, but it is still more omniscient than human players ever have the ability to be, without having to fiddle with the camera etc.
The problem, as they stated in the QnA is that the image processing would take much more hardware and computing power, this increasing both cost and training time, as you would not be able to run games as quickly anymore.