Both absolute and relative timing have to be handled. And relative since specific salient action...
Plus the real reward is very sparse. Say, crippling mineral production early may or may not snowball. Likewise being a unit or two up...
On the specific issue of encoding time-dependent behaviors in models, I think it is related to a broader issue that shows up in many application areas. To me the critical factor is that these models are ruthlessly good at exploiting local dependencies and totally forgetting long-term global dependencies or respecting required structure in control/generation.
This basically means it is very difficult to train long-term, time dependent behavior without tricks (early/mid/late game models, extensive handcrafting of the inputs, or using high level "macro actions"). Indeed, FAIR's recent mini-RTS engine ELF directly gives macro actions, in part to look closer at how well global strategies are really handled and remove one factor of complexity [0].
Gabriel's PhD thesis was entirely on Bayesian models for RTS AI, applied to SC:BW [1], so I am sure he is well aware of the "classic/rules based" approaches for this.
[0] https://code.facebook.com/posts/132985767285406/introducing-...
[1] http://emotion.inrialpes.fr/people/synnaeve/phdthesis/phdthe...
As to the sparsity of reward, I'm not sure this is such a big problem. Once the AI learns that e.g. 'resources are good', it can then learn how to optimize resource production. You could even give the process a head start by learning a function of time+various resources+assorted features to win rate from human games to use as the reward function.
Yes resources are good, but how do you know when to expand?
Judging from opponents movements, you can tell if they're turtling, going for some cheese strat, or doing some build where they may not be able to respond to a aggressive expansion.
Of course if you choose wrong, you lost the game.
With hundreds of unit x objects, jungling, roshan, cd and pick + ban, you can actually get at the sc level of complexity.
So, SC is still a much more complex space. DoTA has non player bots, but they are similar to SC buildings and follow very simple rules.
In dota combinations open the road for brand new moves. Some items tp, some regenerates, some cancel buff, some critics, some cleave, some cut trees, some slow, some stun, some give visions, etc.
Now in SC, you have 3, 4 main builds for a given match up. You see the building, and you know where this is going.
In Dota, depending of the 10 heroes, current money and objects combination, and player skills, you may expect one build or another.
Also, a zergling or 10 zergling is pretty much the same the same to consider from the behavior point of view. The number doesn't matter that much, only the intensity of the effect. And a gling will alway do the same thing. Move. Attach. Burrow.
The same unit in Dota can have a completly different role depending of the context.
My guess is that an AI would give you a much bigger advantage in SC because they can make more APM than a human, strat or not, while on Dota at high level strat is more important on the long run.
For example larva are one of Zerg's most valuable resource and there are several ways of attacking that resource by killing units or simply forcing them to go more defensive.
While that all can be extrapolated from current state I think starcraft is much easier to go for immediate gains by destroying more supply/resource value of units and extrapolate from there.
I noticed this building in this position at this time and I haven't been attacked by X unit yet, so he's probably doing strategy Y. I better skip some unrelated building I was going to make, so I can have an extra unit Z in case he's doing that strategy. Then I'll place the units at a particular spot to try to trap him because that unit will be vulnerable in this other spot so he's unlikely to move through that spot.
It wouldn't be a suprise if some research team could put out a bot achieving superhuman victories purely by out-microing an opponent with minimal strategic choices.
The chess equivalent would be letting Deep Blue take 10 years to evaluate each move; it's not a very interesting system anymore since it isn't playing under normal rules (~90 minutes per turn).
Any "real" SC AI will have limitations on input, say 300 actions per minute. It'd be pretty interesting to see how few actions per minute an AI could use to defeat the top human players.
Even worse and less interesting - it's a bit like allowing the computer to move two pawns in each turn.
Yeah they did pretty much that. But the problem is it's a very brute-force approach and violates some rules of the game.
They jam thousands of commands per second into the game, and give each unit its own rudimentary AI. The units basically just dance at maximum range, magically dodge hits, etc.
If they limit it to 600 actions per minute (10 keystrokes hitting the keyboard every second - still beyond the human mind but beyond human fingers) it becomes a much harder AI problem.
In the case of certain unit matchups, say, zergling versus vulture, the vulture should be able to kill an infinite number of zerglings given that it is microed correctly. However, despite the zergling being useless against a vulture on paper, In a human game you just don't have enough time to babysit your vultures with everything else going on so you end up seeing zerglings being used against vultures somewhat cost effectively even at professional levels.
While it certainly isn't fair to play against, it does have a certain elegance[1].
There's also the problem that even if it's AI vs AI, the races and units are balanced around reaction times of humans.
According to the player interviews and Reddit discussion threads, the "break" you are talking about was more like being really unpredictable therefore finding a play style that the AI had never encountered.
The players were flailing to find a way to defend against the AI that is learning quicker over time than they are.