OpenAI Five Benchmark
blog.openai.com
blog.openai.com
The introduction of the other heroes also comes as a surprise, I wouldn't have expected them to have the ai utilizing new abilities. They don't mentioned how they are picked, other than the ai having a random draft of them (does the ai pick their composition?)
This new constraint is interesting. The Super Smash Bros. Melee AI paper noted that they had to keep the reaction times to superhuman levels in order for the model to converge (albeit DOTA is a bit different from Melee): https://arxiv.org/abs/1702.06230
In fact, for a single step with no policy lag, it's equivalent to a standard policy gradient update.
DeepMind was also able to train a CTF agent with human-level reaction time: https://deepmind.com/blog/capture-the-flag/
I suspect the difference that allows you to train with reaction time is an RNN or compensating for the lag some other way. I'm testing that out right now with my own SSBM bot: https://www.twitch.tv/vomjom
Exactly. Which is why it's so surprising that it did anyway despite that and discount rates which don't give any value past a minute or so.
> DeepMind was also able to train a CTF agent with human-level reaction time: https://deepmind.com/blog/capture-the-flag/
Note that the CTF agent is way more complex, featuring multilevel RL and evolutionary losses, and even DNC in the agents.
The linked commenters thought that getting to "real dota" (more than 100 heroes, captains mode instead of random, ...) would take another year. So I don't think it's fair to make that statement.
Edit: Don't get me wrong, I think the improvements are very nice, but pointing to people saying "these people thought we would need a year, we did it in under a month!" is not what you should do if you didn't actually do what the linked people stated.
Courier isn't that important. It's being phased out across the recant patches. And there is a popular Dota mod with 5 fast invulnerable couriers.
Yes there is Turbo, no it's not comparable to regular gameplay.
So maybe playing without a courier at all would be more representative of the pub experience ;)
So when next time Elon tweeted: "OpenAI beats the human players in 5v5"
You know that the game is not broken by AI yet (not like Go, which is indeed broken by AI).
They'll still have a match against top pros at the International in late August.
My prediction is that we're very very far away from AI that can beat the top teams in a 5v5. Amateur teams can easily be beat simply on the strength of the mechanics (which are very very strong on the AI, beating even pros), but the strategy and coordination of the top teams are out of this world.
Right now they are so many key parts of the game: Illusions, Summons, Bottle, Courier, and most of the heroes. The 18 they have chosen are all fairly straightforward and make drafting simple. I want to see an AI playing Huskar, IO, and Natures Prophet. Better, I want to see an AI that can draft and ban.
The draft isn't everything and it's possible that a sufficiently talented AI could always lose the draft and still win the game, but that would be a pretty boring outcome from the perspective of contributing to AI knowledge (just like it's possible, though unlikely in Dota, that sufficiently good micro could overwhelm any disadvantage in strategy and tacitcs if the AI can play at 2000 APM: it would "win", but only in a very boring sense)
https://blog.openai.com/openai-five/
"We’re still fixing bugs. The chart shows a training run of the code that defeated amateur players, compared to a version where we simply fixed a number of bugs, ..."
Looking at the chart and the fact that they are confident enough to lift several restrictions, I'd bet on OpenAI Five winning against at least some of the professional teams at The International. It's even possible they will beat most teams there.
The output is a trained neural network!
We dump state from the bot API each tick and send it over GRPC to a Python agent, which formats the state into a tuple of Numpy arrays. That Numpy array is passed into 5 neural networks (one per agent), each of which returns a tuple of Numpy arrays. Each tuple is decoded into a semantic action, which is then returned to the game via GRPC.
- I saw no mention of CNNs, is it true CNNs are not used even for the 8x8 terrain grid input?
- do you have any comments about rapid+PPO vs say impala+vtrace? Would the ability to use more off-policy data be very helpful here?
- any comments on how you selected the reward constants?
- was the teamwork/tau something your team came up with or was this a known approach?
- the attention keys are most interesting, can you comment on why they dont flow through the lstm? Does it make it easier for the network to quickly change unit attention or some other reason?
- any comment of the choice of single-layer LSTM vs multilayer ostensibly for operating on longer timescales?
- does this result mean that HRL is less critical than some people thought?
- any comment on magnitude of compute, like in the post from may?
Thank you for sharing your fascinating work!
https://twitter.com/eternalenvy1991/status/10196414446030520...
- Clearly defined goal
Overall, games are a good playground to test ideas and verify assumptions. The next step to transfer this type of knowledge to real world problems would be to build a simulator, train on it using ungodly amounts of computing resources, and then fine-tune the final model on the real world thing. This has been done for robot control tasks in the past. But first, you have to develop and prove that the base learning algorithm works -- and games are nice for that.
This here is also a good showcase of collaboration learned by RL agents, and beating pro teams in an esport where prize pools range in the millions of dollars is an amazing way to convince people.
And we, humans (and animals) have a huge environment with billions of agents and millions of years of evolution behind us which allows us to come preloaded with good instincts, they are trying to replicate this process in a few months.
How does a deep learning algorithm coordinate between 5 heros? I assume it's not 5 bots communicating over some channel but one bot acting on 5 heros?
"OpenAI Five does not contain an explicit communication channel between the heroes’ neural networks. Teamwork is controlled by a hyperparameter we dubbed “team spirit”"