DeepMind and Blizzard to release StarCraft II as an AI research environment
deepmind.com
deepmind.com
DeepMind’s last triumph (beating the best human Go players with AlphaGo) is impressive, but Go is a great fit for neural networks as they stand today; it’s stateless, so you can fully evaluate your position based on the state of the board at a given turn.
That’s not a good fit for most real-world problems, where you have to remember something that happened in the past. E.g. the ‘fog of war’ in a strategy game.
This is a big open problem in AI research right now. I saw a post around the time of AlphaGo challenging DeepMind to tackle StarCraft next, so it is very cool that they have gone in this direction.
When Google’s AI can beat a human at StarCraft, it’s time to be very afraid.
I don't actually agree with this. Unlike Go, StarCraft is not only a game of strategy; "micro" (micro-managed tactics, basically) also plays a big role. An AI is going to be able to issue a LOT more commands per second than even the most skilled humans, giving them a natural tactical advantage.
Strategy is more difficult for an AI, but not the only deciding factor in StarCraft.
More importantly, in my opinion, I'm not aware of any other AIs that have been competitive at the StarCraft problem using a Neural-Net-based implementation; from talking to PhDs in the field (but not from direct technical experience, so take this with a pinch of salt), my impression is that state-of-the-art RNNs just don't have the ability to encode the history of the game-world in a way that would enable them to be competitive. So it will require significant advances to the theoretical models to be able to compete using the sort of implementation that DeepMind will build.
I tried to make a blink stalker rush that'd be the rough equivalent of the build shown in the post liked above, just with the Protoss race in the game. The program would tell me to do things like make two extractors but put fewer than necessary probes on them, pull probes to extractors after X trips or at X time (can't remember which), things like that.
1. http://lbrandy.com/blog/2010/11/using-genetic-algorithms-to-...
I had to look it up again to be sure, but 2 on minerals 3 on gas haha: http://www.teamliquid.net/forum/sc2-strategy/140055-scientif...
Pros actually do both of those things.
Otherwise the AI is eventually just going to absolutely destroy humans with micro, which is not remotely interesting. We've known transistors were faster than fingers for ages.
Also, I'm not sure how they plan to limit APM, but if it's done in the naive way - the AI simply isn't allowed to surpass an APM of X, there will be plenty of optimizing it can do in terms of spending those actions more optimally than a human could (in some cases by performing fewer but more difficult actions to accomplish a task), as well as by remaining at peak APM continuously throughout an entire match. (Something that is a hallmark of great players, but like map awareness and macro timing, no human is able to keep up their peak speed indefinitely without any lapses.)
Whenever the harder action is better than the simpler one, the AI would gain an intrinsic advantage even though it's limited to the same number of APM.
If you did NOT limit APM, it wouldn't even be a contest. The AI could simply mass first tier units and crush any human opponent with micro.
The goal is to build an AI that is at least as good as pro players (which would involve building models and technology capable of taking complex decisions etc.). The goal was never to implement something that is strictly equal to a human.
You cannot blame the AI for having a better memory and perfect timing, or next thing you know you'll ask them to limit the RAM to something comparable to the memory in a human brain and use older processors to restrict the unfair advantage current computation clusters have.
Regarding your last point: even without limiting APM, there is no AI that can currently compete with decent players . If it was as easy as you say, people would've done that with Broodwar which has open APIs. The fact is that micro alone cannot win a StarCraft game. Perfectly micro-ed zerglings cannot do much against a terran turtling behind a closed wall and shooting with marines, or a protoss with immortals and adepts.
Could be interesting to see if it chooses different strategies based on that. What actions will optimize performance if I can only do 50 a minute
And it worked for him, too.
If you could issue as many orders as you wanted, what would the optimal strategy be?
Here's what you can do with infinite APM: https://www.youtube.com/watch?v=IKVFZ28ybQs
It's just broken. It's not even the same game.
(The implementation in the video likely just cheats and gets the information from the engine, but at least the first case should be deterministic enough)
Winning against AI in a sufficiently complex game looks very different than winning against a human.
A lot of units would be rendered useless because it's balanced for humans. That means the units they counter would be a lot stronger. I'm afraid it would probably just be a race to get the fastest rush.
Here's the type of micro pro players can do: https://www.youtube.com/watch?v=xrAlhk98WxE±
They do say that they will limit the number of actions taken by the AI to focus on building a smart and strategic AI instead of learning gimmicky micro moves, but if one thing is sure, it's that even the best micro cannot win a StarCraft game if the macro is not on point, and bad choices are made (good luck micro-ing siege tanks against a swarm host!)
Meta is still pretty fast, here's a SC1 AI finals: https://www.youtube.com/watch?v=LjSXj4cb_Yo
There's would not be one optimal strategy, but many optimal strategies depending on the matchup between units. By taking advantages in differences in movement speed, attack range, attack projectile speed, and attack cooldown, a weaker group of units can potentially beat a stronger group.
This is an older video that demonstrates some examples of winning with good micro: https://youtu.be/CdSKD3LRHV8?t=11. An AI knowledgeable of and capable of executing such tactics would be pretty strong.
One example is a sibling post's video of 100 zerglings against 20 siege tanks. For humans, the zerglings are capable of destroying only a handful of tanks before being killed of. For an AI with infinite APM, they can move all but an individual targeted zergling out of the way of the tank's splash damage. This allows them to take out all of the siege tanks with only moderate losses. Many such matchups will be tossed on their heads, and it's likely that one of them is totally game-breaking.
They regenerate, so the AI just has to keep them alive and do damage to win.
Like a real army :)
just to illustrate a point: on a sufficiently large map, i could have four times as many workers by the time your first set reaches my spawn.
if i'm playing Protoss against a Terran AI, i have an added advantage of worker regen without having to stop attacking.
Also, I'm pretty sure this is not something that DeepMind would do to "solve intelligence".
How much more complexity is added by resource mgmt vs just "legal play possibilities" as now you have "this move isnt legal because you dont have the requisite resource, thus you need to depend on completing those moves first"
Are there any efforts applying deepmind to, say, commodities markets? Or something along those lines? Supply chains for complex hardware manufacturing logistics?
Robot assembly of complex parts in an assembly line?
Sure, AI's playing games is cool...
What I would like to see is an AI playing "Farm Simulator" to manage a vast swatch of farm equipment making and picking food over a vast expanse - with little human interaction.
* Auto dock
* Auto Harvest
* Auto Refuel
etc....
That would go well with the concept of Seed Factory, which is a factory that could reproduce itself, from raw materials. If we can build a self replicating device, we will reach another kind of singularity. At the moment, the whole human race is such a self replicating system, but if we can take humans out of the loop we could ship it to an undeveloped area (another planet?) and have it create a city and industry for us.
Another application would be to ensure self reliance for humans in the post automation world. A self replicating factory, capable of solving the whole logistics puzzle of building repairing itself, as well as supporting its owners, would ensure that its owners don't depend on UBI or jobs.
Such a factory would initially need some industrial machines and robots, probably at minimum a 3D printer, some raw materials and high end pieces such as CPUs. Then it would gradually print and assemble more equipment.
It would eventually bootstrap like a compiler, being able to "compile" an identical factory from raw materials.
Keeping count of each object, its source parts, construction waste, energy spent in building it, processes and time they could "compile" an ecosystem of devices that carefully work together, recycling everything that is wasted in the process. As the capabilities of the factory increase, the process could be "recompiled" to take advantage of the new capacity.
The place of AI would be to optimize the process. Instead of a hand-made industrial compiler, it could search for a way to combine its resources in more masterful ways. It can keep track of more things than human planners (project managers?) can. It could potentially assess the needs and organize the industrial production for a whole country in a more efficient way than capitalism or communism could.
The beauty of a seed factory is in its initial price - under 5% of the price of a factory built from buying equipment, because it only needs the seed of the factory and raw materials. The factory would grow like an organism, expanding itself. Most of the factory would be self made. Just like a cell, or an animal.
[Wiki Book] https://en.wikibooks.org/wiki/Seed_Factories
There is a massive advantage to seeing everything all the time and being able to arrange and organise everything at the same time.
One AI won simply because it abused the fact it could spam monks and use the monks in a manner no human could ever achieve
Don't get me wrong; I'd prefer DeepMind's efforts be focused on a huge and realistic economics simulator but there are simply no such games (not ones that are good and/or realistic anyway).
IMO StarCraft 2 is the best economy and war simulator these days so that's why DeepMind chose it.
Having an AI that wins because it is faster isn't very interesting.
And they are really bad.
You could solve this by coming up with some sort of context-switch effort function but it starts to feel forced.
Or at least, once you start calibrating the parameters of the problem around human limitations, it sort of feels weird to talk about the computer being better or worse than the human.
Doesn't seem too forced to me IMO. Humans have to expend more effort to context switch, so if the goal is to put humans and computers on a level playing field that's something you have to take into account.
They address that in the article. The context they're talking about is a bot beating a human with human-like dexterity limits in place.
Seems like they should probably tackle a game that involves strategy without micro as a stepping stone to StarCraft. Are there any games with good AIs like that already?
It doesn't matter how well you split your zerglings or move your marines, when there are tanks busting through your front door.
But I totally agree, long-term dependencies are an difficult problem. The DNC[0] by DeepMind is in my opinion a very important step in the correct direction to tackle this.
[0] https://deepmind.com/blog/differentiable-neural-computers/
I can't imagine that this is where DNCs come in unless you want to needlessly complicate the matter (and create a network that is much more expensive to train).
From the article it seems the answer is: very little. It is allowed to use only the pixels on the screen. (During play, not training apparently.) It will even have to spend actions on moving the screen/window to inspect things. Very cool stuff :)
There's already been excellent progress in poker AI based on traditional game theory, see http://poker.cs.ualberta.ca/.
For instance, as any protoss player knows this all too well, a Zealot placed a mere pixel away from where it should could allow a stream of zerglings in, completely throwing any game plan you had up to that point out the window and enter crisis management mode. Usually this is a "mistake", which a computer may never make, but it needs to learn the relevance of such pixel-sized mistakes, which is unlikely to have as much of an impact in a poker game or Go game as it would in StarCraft.
I agree with your point, but you need to look no further than the game that AlphaGo lost against Lee Sedol.
It committed a very bizarre mistake that made no sense gamewise.
I think that there are two distinct aspects of the incomplete information that are significant; first, ignoring temporality, you have the fog of war, so you can't see the whole board. This is probably easier to address in a RNN, since you can play quite well as an amnesiac that just reacts to things that are currently visible. But you need to scout less if you have a memory of what's out there, so the amnesiac won't be able to play optimally.
Then there's the temporal aspect. The set of previous states of the game is not stored in the game and made available to the player, and so to play optimally you have to have a memory. This is where new techniques will be necessary.
These are separate problems, I think, so it will be interesting to see if DeepMind can make progress without reaching human performance on the second part.
https://www.youtube.com/watch?v=IKVFZ28ybQs
If the US doesn't do this, someone else probably will.
We're going to see an AI-based blitzkrieg attack in our lifetimes.
Once you can launch an invasion without losing any soldiers, war seems like a much less dangerous proposition.
Yes, but how can you stop everyone else from doing the same, if the AI part will be trivial to replicate?
That's not always true: https://en.wikipedia.org/wiki/Ko_fight
"[ neural networks ] are not a good fit for most real-world problems, where you have to remember something that happened in the past."
This results in a very low overall competency, when it comes to winning games against highly skilled opponents.
It will be a challenge for researchers to improve performance in these areas. The problem space is very complex -- it makes solving Go look like a cake walk in comparison.
Im curious if a startup can be built from this.
Making an AI for game isn't about making a good AI. It is about making an AI that loses in a convincing manner. This is especially try for games like Starcraft (RTS and strategy games in general. For FPS games it really isn't so important as you can just increase or decrease the accuracy and hp of the enemies)
One commonality, it seems, is that players want to get better at the game. Perhaps, an AI that TEACHES the user to improve their play, and reduce errors would be an easier problem. So far we mostly do this using heuristics of gradually increasing difficulty, but it should really depend on what the player is weakest at... and what the user can improve at fastest! That is something nicely measurable, which you might use to train an AI.
Then, if/when people play against other human players they are as strong as they can be... often an FPS goal.
That's how MOOCs should function, too.
This is actually a way more interesting problem to solve, for the gaming industry.
Beating games is cool for industries where gaming is not the final problem (autonomous cars, various decision making, etc). But if AIs are to be exploited by gaming industry, it needs to learn things like : "is the user bored?", "what usually triggers new interest?", "how should difficulty be adjusted given current way the user is playing?", etc.
Wouldn't you be able to do mutual training that way? By mutating the AI's "strategy" only after losses, the player would continuously need to figure out a weakness in the AI's game-plan. And the AI tries to evolve its strategy every time it loses.
And a player in the top 20 percent could beat any AI that has ever been created.
AIs are really in an abysmal state for RTSes.
See "The Marginal Advantage," an essay from the famous Starcraft player Day9: http://www.teamliquid.net/forum/brood-war/64514-competitive-...
It's also worth noting that AlphaGo optimizes its perceived chance of winning the game, rather than its score. It "gets complacent" and mostly plays very conservatively the moment it's ahead on points, because it doesn't care about margins, only about win/loss.
I meant to say that strategy trumps mechanics. Mechanical proficiency is necessary but not sufficient.
Additionally, humans can use economic reasoning (player built 3 barracks to they can't be building x, y and z). This can lead to AI being excessively safe (economically inefficient).
And the last line you said is also true to make things even more difficult. :)
First person shooters can have super accurate AIs that are really frustrating to play against. Try beating the last "boss" in quake 3 on nightmare difficulty. A game like Starcraft relies much more on decision making that raw 3d math calculations and that makes making a decent AI pretty difficult without giving it advantages like extra resources or vision. That said, even though a game like Quake is easier to write an AI for, top players can still be top level AIs with superior reasoning about using the map and timing power up spawns and such.
The Worms series is also notorious for this. Their AI targeting seems to flip between "clueless" and "godlike", with very little middle ground. E.g. https://youtu.be/gqPITW04vRQ?t=1m20s
Starcraft however has all kinds of minor inefficiencies you can build in. You could inconsistently build workers, you could occasionally supply cap, your army comps can vary in sophistication level, etc.
It is actually about making a good AI. Civ 6 or SC2, in both cases they simply "cheat" by providing multipliers to the AI in order to solve its limitations. The basic function is largely the same. Providing the AI with cheat codes is not "good" AI.
A good AI design would remove the need for such arbitrary advantages that are out of context for human players and leads to weird issues.
Where the AIs tend to be very bad is in long term thinking and overall strategy.
Blizzard's AI in Starcraft generally plays like a slightly more aggressive but novice human. It's more fun than the TA AI, but not terribly strong once you understand the game mechanics.
What wrong with a strength option? I'd include an unbeatable entry in any case.
The same is sort of true in RTS games. The new rerelease of age of empires 2 has vastly improved AI that was developed by modders. Lots of players thought it was just cheating.
But in general it has been very well received. I think players can get used to good AI. They've just been trained to expect stupid AI that cheats.
Still, AI in games is mostly a trade-off. It usually can't take too much resources (because it usually has to run on the same machine the human is playing on), it has to be believable and fun to play against. This usually rules out too fancy algorithms and approaches with dubious returns.
Also, out of my experience with many different fields of programming, game programmers seem to want to write more of the stuff themselves than most. This would be another huge subject to go into, but I think the biggest reason is that historically you couldn't ever really punt on performance in games, and you had to tweak everything to your problem. If your game ran slowly or looked bad, people just wouldn't buy it.
Doesn't quite apply to a mario style game
Basically set up a queue of player and what NPC character to play and when a user gets to that part of the game where NPC needs to load, you jump in and play. When the NPC dies you the game (scene) and get put back in the queue of playable NPCs.
Unfortunately, humans will find advantages that shouldn't be used in any sort of game that tries to build a believable world.
That isn't to say it's a bad idea. You'd just have to design the game with human enemies in mind. I was very interested in The Crossing [2], which incorporated that of mechanic. Unfortunately, it was canceled. :(
[1] Not a made-up problem. https://youtu.be/UfZFHdywYLA
How would you scale for games with different inputs and outputs?
it's necessary anyway to prevent it from having godlike micro.
Allowing researchers to build AIs that operate on either game state or visual data is a great step, IMO. Being able to limit actions-per-minute is also very wise. The excellent BWAPI library for Starcraft Broodwar that is referenced (https://github.com/bwapi/bwapi) provides game state - and was presumably used by Facebook to do their research earlier this year (http://arxiv.org/abs/1609.02993). For mine, the significant new challenges here not present in Go are the partial observability of the game and the limited time window in which decisions need to be made. Even at 24 frames per second, the AI would only have 40 milliseconds to decide what to do in response to that frame. This is more relevant to online learning architectures.
The open questions here are how freely this will be available - and in what form. Will I need to buy a special version of the game? Clearly there will be some protection or AI detection - to ensure that competitive human play is not ruined either by human-level bots, if they can truly be developed, or by sub-par bots. Starcraft 2 (presumably the latest expansion, Legacy of the Void, will be used here) does not run on Linux, whereas most GPU-based deep learning toolkits are, so having a bridge between the game and AI server may be necessary for some researchers.
Besides being great for AI researchers, this is probably good for Blizzard too, since it will bring more interest to the Starcraft series.
2017 can't come soon enough.
Let's just get through Tuesday and see what happens! ;-)
I'm not sure how familiar people are with StarCraft II, but at the highest levels of the game, where player have mastered the mechanics, it's a mind game fueled by deceit, sophisticated and malleable planning, detecting subtle patterns (having a good gut feeling on what's going on) and on the pro scene knowledge of your opponent's style.
I wouldn't be surprised if an AI comes along that shows this human perception of what it takes to win isn't required at all. I'm sure similar things were said about Go and Chess.
I took deceit to mean tricking your opponent into thinking you were doing one thing when you were actually going to do another. You can't do that in Go or Chess?
What I generally meant was maybe there's some really boring numerical number crunching or some simple mechanical algorithm that solves games like StarCraft but when humans talk about it, they talk about mind games, deceit, intimidation etc. when maybe these things are irrelevant.
Remember that it's easy for an SC2 AI to beat a human if there are no constraints such as APM. So if and when the AI wins, the humans are going to complain that the constraints weren't set tightly enough.
So the contest is going to devolve into a war over the constraints to use, just like the endless bickering between Protoss Terran and Zerg about balancing the units.
The best thing DeepMind will be able to say is that they were able to beat Flash (or whoeever) with only 150 APM. But Flash will complain that it wasnt fair because DeepMind was able to be so efficient with those 150 APM so that really it should be limited to 100.
I suppose once the APM cap is well below what humans do, then the humans will have to admit a sort of defeat.
Hopefully there will be a ladder running where humans can play against the AIs under development and get a feeling for how they play. Will be really fun & interesting!
The same is true of Go.
It's different yet similar.
I meant it in the other direction; what if top players rely extensively on deceiving each other, but an AI is better able to predict its opponent's strategy than a human an isn't fooled? Again, I'm not saying this is necessarily the case, but I don't think it's implausible
[1] http://bloomreach.com/2014/12/centaur-chess-brings-best-huma...
The closest thing I can think of is add-ons for MMOs, but there the game developers specifically try to prevent add-ons that are too performance enhancing instead of making the coding an integral part of the game.
Ontanón, Santiago, et al. "A survey of real-time strategy game ai research and competition in starcraft." IEEE Transactions on Computational Intelligence and AI in games 5.4 (2013): 293-311.
http://webdocs.cs.ualberta.ca/~cdavid/pdf/starcraft_survey.p...
The top contestants were employing strategies that required APM an order of magnitude higher than what players could do. For example, in brood war, SCVs can repair any Terran building or vehicle but it is not worth the effort to repair goliaths or tanks for a player.
However, the AI could easily manage repairing it's tanks which makes the Terran army much more cost effective.
In SCII there is auto repair and other improved UX that probably limit the ability of novel strategies like repairing vehicles.
An AI limited to even 200 decision-driven actions per minute would probably have a significant advantage over a "500APM" player.
Unlimited APM makes AI much less interesting from the perspective of game theory.
Think of it like a robotic boxer, except instead of an android, they build two 30 foot long walls of spring-loaded boxing gloves that close in on the human boxer. Yes, the robot punches the guy a lot, so the problem seems "solved" but it isn't really.
The point is to build ai that compete with humans using human strategies. Not to cheese with borked game mechanic exploits.
Kudos to both Blizzard and DeepMind. Anticipating a lot of fun with this. StarCraft 2 could indeed become the AI pedagogy standard.
Inter-agent communication, especially in a machine learning context, is an unsolved problem, exploring it in such a scenario would be nice.
Agent-human communication is probably as hard, and also largely unsolved, although there is a push right now for dialog systems to achieve human level performance.
this brings up another interesting topic all together- which is how inefficient human language is at communicating complex ideas, quickly, with people you don't know.
that is- just like how AI has an obvious APM advantage, it also has an obvious communication advantage (raw serialization and deserialization. there is no room for ambiguity or misinterpretation).
basically, all information known by one agent will be known by all other agents in near-realtime. something humans can't do yet.
I hope some of the advances in SC2 AI can be integrated into the in-game AI. e.g: a trained neural network that plays better than the "hard" AI, but can run on a consumer box and not on a massive cluster.
I think the worst possible outcome for society would be if we ended up with capable AI but with the algorithms only accessible for a handful "AI-as-a-service" companies.
The second concern is understandability of the algorithms: from what I've read, it's already hard to deduce "why" an ANN in a traditional problem behaved like it did. An AI that is capable of coping with dynamic, unpredictable situations like an SC game (with only pixels as input) is impressive but it seems less useful to me if it's not understood how that is done.
This is awesome. I've only ever reached the Platinum league in Starcraft II (1v1), but I'd almost feel more driven to create bots to (hopefully) surpass that skill level, than actually playing the game.
I sometimes play strategy games and I always find the AI disappointing. Any game with a great AI would be my favorite for years. Heck, I would even pay a few dozen cents/hour to be able to compete against a proper AI.
But I've never done it because of the risk of bans. I'm glad that Blizzard has opened it up for people to experiment with this. I wonder how it will interact with any sort of anti-cheat systems in place, etc.
I personally am willing to pay $40 for an environment to test my ai on for a couple months.
Starcraft isn't troops on the battlefield, in the same way that Risk, the boardgame isn't troops on the battlefield.
Just because the 1s and 0s or game pieces represent "troops" to humans, doesn't mean that the underlying game mechanics have anything at all to do with a real war.
For all we know, the problems solved by a monopoly AI, are more applicable to a real war than those solved by a Risk or starcraft AI.
Similarly, algorithms used to detect cancer in medical imaging can be used to do automatic targeting on killer drones. Should we stop doing research altogether?
Legends trace the origin of the game to the mythical Chinese emperor Yao (2337–2258 BC), who was said to have had his counselor Shun design it for his unruly son, Danzhu, to favorably influence him.[64] Other theories suggest that the game was derived from Chinese tribal warlords and generals, who used pieces of stone to map out attacking positions.[65][66]
Of course, I doubt that an AI commander would replace a human making the actual decisions in the foreseeable future. It would just be another intelligence tool for military command. From this point of view you are essentially arguing for our military to be ignorant, I don't think that is a good bet.
Or maybe the answer is never, other companies are supposed to do the hard work? We only play games?
https://deepmind.com/blog/putting-patients-heart-deepmind-he... https://deepmind.com/blog/applying-machine-learning-radiothe... https://deepmind.com/blog/deepmind-ai-reduces-google-data-ce...
Projects like these may have a "fun" goal, but while reaching towards that objective, a lot of new techniques and theory is developed, which then gets used in real products.
Systems like this are much easier to use as a playground and test in than the real world. But as they mention, the messiness and stateful nature of games like SC2 will bring a lot of innovation which will be very relevant to the real world.
The reinforcement learning they work on is extremely general and can be applied to many other types of problems. Particularly robot control, which is sort of the same sort of task as a video game.
Who do you mean by "the best humans"? I've never heard of top-tier progamers such as Flash, Jaedong, Bisu, etc. losing to AI. In fact, I doubt any B-teamer or above would lose to an AI in a Bo3.
Why not give it lots of data to solve real problems? Training it on useless games will have no benefit.
Starcraft is a really fun game, and I think it's enough to engage kids a little more than something like Minecraft where there's plenty of room for some cool ML hacking, but not enough stimulation from it. Instead of just seeing blocks here or there or whatever, starcraft has hard goals that will force them to use critical thinking skills, their knowledge of the game, their own personal strategic insights, and the ML skills they accrue.
So exciting! Love the feature layer idea also, well done!
DeepMind really chose well. SC2 has to be the most demanding game nowadays in terms of strategy and execution.