Deep Reinforcement Learning to Play StarCraft
arxiv.org
arxiv.org
So perhaps it's worth pointing out, that this paper specifically addresses a sub-problem of Starcraft play, micromanagement ('micro') [1]
The game engine runs at 24 frames per second. (As an aside, 'frames' in this context likely does not map to physical FPS of the display).
>We ran all the following experiments with a skip_frames of 9 (meaning that we take about 2.6 actions per unit per second).
The research team found that attempting to move at a superhuman pace (eg one action every frame), resulted in a subpar performance and hyper-parameterization indicated 2.6 to be an ideal action per second.
In context, this translates to an APM of 156. Or, roughly half that of professional Korean e-athletes. [2]
[1] https://en.wikipedia.org/wiki/Micromanagement_(gameplay)
Moving at extremely fine-grained timesteps can make learning much more difficult, because now a reward arrives millions of timesteps delayed rather than hundreds or thousands. It's like trying to teach a NN to compose piano music by starting down at the 1ms raw audio level. This is part of why audio synthesis was so difficult up until recently with DeepMind's WaveNet. In theory, being able to move every frame should enable extremely superhuman performance, but in practice, you can't learn your way there. So often people will chunk data to make it easier to learn the higher-level concepts: operate on words, rather than characters, for example.
I guess that's inelegant when a deep network already has its own concept of fine-grained versus coarse-grained layers, and should be able to do this on its own with the right training method.
It would then seem to suggest, that forces with more than say, 5 units (780 average sustained APM) and above, would likely be getting into super human territory.
You can also hotkey individuals or groups, 1-8. Again just rolling through those will get you a fair amount of apm.
That said, 780 is peak for world class players. You probably can't do that, ever, because by the time you put in the years of effort, your reflexes will have decayed.
On the other hand, brood war was a subtle game. APM is not the be all end all of winning. Getting proficient with a few aspects of the game can make you pretty good pretty quickly. Heck, just looking around the map like i mentioned above will help. Some people are amazing controlling a few units like a surgeon. Those units are practically unstoppable. But they miss the big picture. Scanning the map lets you distract, delay and ultimately minimize the value of those unstoppable units.
[edit] Ah, $21k, not $31k. Thanks. [edit] And yes, it was recently completed and isn't actually ongoing. Not sure what I was smoking when I made so many false statements.
[edit] Maybe I was thinking the tournament starting October 28th, http://wiki.teamliquid.net/starcraft/VANT36.5_National_Starl... for $33k total prize pool.
There isn't something similar available for SC2 due to a mix of technical and nontechnical issues:
https://github.com/bwapi/bwapi/wiki/FAQ#will-there-be-an-api...
I haven't played for a while, but usually there are a few basic 'builds', which are essentially memorized openings-- and there are 3 types of openings -- Macro builds where you focus on building an economy, while sacrificing military resources for a long-game, 'all-in' builds, which sacrifice your economy to build an early military advantage and win within the first few minutes, and various mid-range builds that try to do a little bit of both.
An all in is largely just down to micro and execution and it either wins or it doesn't, but the other two types of builds have a large strategic element-- for example, you need to scout to check if your opponent is all-in-ing, you can do harassment to distract your opponent from implementing his strategy by interrupting his economy, and then there's planning for the end game, building defenses, and the whole question of what you do if your original plan fails for one reason or another. There's a lot of thinking involved on multiple levels simultaneously, both spacial and temporal.
Here is a video of one of the best current StarCraft bots losing to an D-rank (low skill) human player. The bot's APM is ~5500 while the human's is ~200. https://www.youtube.com/watch?v=ztNYOnx_YQo
The fact is, no AI has ever beaten even an amateur player in a tournament. Even with great micro, if your play is too predictable then the human will learn it and exploit it.
I, for one, am very excited to see the development of new StarCraft AIs. And especially SC2 AIs so that it can challenge the current world champions.
What kind of things did he prepare? It wasn't reflexes, it was strategy. What kind of strategies was he preparing? He watched his opponent's past games, and came up with some build orders of course, but in this case, the primary strategy he came up with was an army composition hoping to counter what his opponent had been doing recently. When the opponent had the proper counter to that strategy, he won the rest of the games easily.
Seems like there's more potential for useful AGI techniques in that direction.
It would be interesting though. How would a program that has PERFECT micro fare against a professional Starcraft player. Would a program reliably figure out how to kill 10 banelings with a few marines and a medivac by using the fact you can micro them to be able to do it without taking losses? Even if it could, would it know WHEN its even worth it to do so?
The "Hard" part of StarCraft is that it is a huge Rock Paper Scissors game with only the information you fight for. You have to be able to piece together a picture of your opponents actions and forces from small cues.
Here is "Automaton 2000" controlling 20 marines vs 40 banelings, without losing a single unit.
https://youtu.be/DXUOWXidcY0?t=52
Pretty cool.
It's enough to make you scared for the future of humanity.
As you said, the glory of StarCraft is it's strategic level information game. Will be interesting to see what comes out of attempting to learn that.
http://webdocs.cs.ualberta.ca/~cdavid/starcraftaicomp/report...
The two papers in RTS techniques sections are a must read for an idea of what problems it poses along with results of prior attempts. The ability of human pro's to detect AI patterns and defeat them with bluffs is pretty consistent. StarCraft, like Poker, involves lots of psychological analyses and ploys.
Even if Google or Facebook make one, I still think of humans as superior until it can learn how to beat them with mere dozens to hundreds of games rather than what was fed into AlphaGo. That wasn't human equivalent or superior so much as approximating the results of nearly all human activity in the space then focusing it against one human. You could call it superhuman but it required tons of activity by brilliant humans. Brilliant humans require little with the champions a lot less than the automated techniques. Lots of self-discovery with limited data. I want to see the AI's pull that off plus keep it going when encountering humans with innovative, never-before-seen strategies. That's when I'll give them credit as useful on barely-defined problems with curveballs like humans.
We're also discussing support for such projects using game replays from HSReplay.net :)