OpenAI Five: Goals and Progress
openai.com
openai.com
Insta-TPing right when an enemy wastes their stun and can't cancel their TP.
Grouping up as 5 at the beginning of the game and pushing into the enemy jungle. Pubs never do this.
The most interesting part is that OpenAI appears to be discovering new knowledge in the dota scene. For example, they always take the ranged barracks first, never the melee. This is exactly the opposite of what the pro scene does. Therefore, the smartest pro team should study what the bot is doing and trust that on average it's a better idea to always focus on the ranged barracks first. After all, if it was a bad idea, they probably wouldn't do that.
The most hilarious part was when OpenAI paused the game, then resumed it. This illustrates that there is still some unexplainable randomness.
Question for OpenAI: Is it more accurate to think of the bots as 5 separate minds, or a single mind controlling 5 heroes?
EDIT: By the way, TI is going on right now! https://www.twitch.tv/dota2ti If you're new to the scene, take a peek. TI is always so high energy -- even if it's hard to follow what's going on, listening to Tobi (the shoutcaster) go nuts during the game is always a highlight.
And of course, /r/dota2 has the best memes anywhere, hands-down. https://www.reddit.com/r/DotA2/
People were astounded that the SF mid was creep blocking at an extremely high level, since wave positioning/management is a very complex behavior/process that doesn't pay off directly right away. The OpenAI team said that they basically taught the SF specifically how to creep block, and they've done this with other behaviors as well, like denying allied creeps, as well as a lot of the meta-level strategic play like warding/vision and item builds/usage.
For example: There are an infinite number of places to place a ward. One way to train the bot is to preselect all possible ward locations, reducing it to 30 or so common ones. Another way is to make an optimization algorithm where the bot focuses on trying to maximize the "strategic vision" (if it's possible to come up with a measure for "strategic") and then let the bot place wards wherever it wants. After hundreds of years of self-play, it should figure out the best place and times.
As I write this out, I think you're probably right. There are too many aspects of the game for a purely-random algorithm to be effective... E.g. item builds. But I'm holding out hope that's just because they haven't figured out a good way to encode all dota items into a distance measure.
AI has proven over and over that humans aren't so special. And humans know how to adapt to the game.
That said, I wish OpenAI would be completely transparent as to what's emergent behavior and what's not. :)
Oh, one last interesting thing: Icefrog is going to roll out a big patch after this TI, just like he always does. I wonder how much of the bots' knowledge will transfer over? Or if they'd be better off training from a clean slate?
Totally agree with you here. This would also help actual DOTA players to analyze what is truly emergent behavior that might open new doors to high level play styles, and what is just weird coaching/niche optimization choices by the OpenAI team.
The bots aren't even playing real Dota, let alone current patch.
OpenAI are here to compete on equal footing. They aren't going to stick with the old patch, I would guess.
It's sort of a "Theseus's paradox" situation.
It's still really fucking good. But it's not dota _yet_
This is a pretty good summary - the main bit to know is reward shaping https://medium.com/@evanthebouncy/understanding-openai-five-...
I disagree. I think AI has shows that we are very special.
I think it would also raise the level of play that we see from pros. That's what happened with chess bots when they got good enough to be the best human players.
This will always exist in dota, but being able to simulate hundreds of years of drafting + games could help reduce this.
Can you explain?
Take OpenAI game 3 as an example. The first two games, OpenAI wiped the floor with the humans and taunted them that they had a >90% chance of winning. The third game, OpenAI was saying the bots had an >80% chance of losing by 5min. The sole difference was the draft.
This is a bigger factor in my opinion. Each team alternates two bans, then two picks, then there is another round of two alternating bans followed by two picks and finally a single ban and pick round for the 5th hero.
Various in-meta heroes are usually "first pick/ban worthy" which means they tend to get picked or banned in the first phase and tend to shape the rest of the draft as teams will build the core of their strategy around the first phase heroes or around countering oppositions first phase heroes.
Another strategy is to avoid "showing your hand" during first phase by first picking strong but generic heroes that can fit into many potential lineups to keep opponent guessing. This leads to a lot of mind games where even commentators don't know what role the hero is going to be played in until the culminating 5th pick when the draft comes together.
Some teams are very good at specific strategies or have certain players exceptionally skilled at individual heroes which necessitate certain first phase bans against them lest they have an advantage.
For instance If a team is known for having a player good at the hero "Wisp" it will often force out a first phase Wisp ban from opponents because it is the kind of hero that when played well can be absolute nightmare to play against.
In some ways I find the draft mini game to be just as interesting as the main game especially in the longer tournaments where you can see new metagames emerging as captains adjust their pick strategies.
(I'd handicap the crossover somewhere between one and two orders of magnitude more expensive than that.)
Whenever my team is pushing high-ground and we're not certain we have the time and strength to completely 100-0 the melee racks, we always try to hit the ranged one first, otherwise we'll just get pushed back off the high ground and all the damage we did will just be regenerated, negating the value of all the resources we expended in the push.
Ideally, we prefer to take the melee if we know we can, since there's more melee creeps in a wave, the buff is worth more, but if we're not sure we'd rather do damage we know won't evaporate in a few minutes.
This is true - but not the whole story. Since ranged creeps do significantly more damage taking the ranged rax means that your lane still pushes a bit and importantly it accumulates ranged creeps as it goes along. So if the lane is left, when it hits their base it will do hella damage to buildings if they don't address it. (Bonus melee supers would last longer and do more base damage over their lifespan - but what you want is damage in the time it takes them to TP).
(also trash tier player, so just my thoughts)
When training starts out, the bots solely focus on their own reward functions. This is so they can learn very basic things like how to move around, how to attack, and so on.
Over time, the combined team reward function gradually gets weighted more and more heavily, so that teamwork is encouraged.
When OpenAI wins against human opponents it seems mostly be because it is so much better at cooperating and can jump the enemy so quickly. That combined with continued ferrying of regen that would not work in a normal game.
The push heavy death ball strategy is pretty much optimal for the high-precision, perfectly coordinated, mechanical fighting skill the bots have. They get a few kills using the mechanical skill and then they group up and press their advantage as hard as they can.
It seems like all the rules/mechanics they are still working on are the more abstract out-of-the-box stuff that evolved to deal with these types of 5-man "teamfight" hero lineups... (warding, game-contextual item builds, courier management, full hero pool with all the "rat" split pushing heroes, etc.)
Neither AI experts or DotA experts have the necessary background to predict how close anything is to any level of future performance. DotA players are probably the worst to ask, because the AI is already stronger than they are at a subset of the game and they have no insight into how quickly that subset could be generalised into other aspects.
OpenAI has a list of something like 7 restrictions. Nobody has any real idea of how quickly these restrictions can be lifted once the OpenAI team has an AI that has mastered the game with those restrictions.
Eg, 5 invulnerable couriers is obviously a huge restriction - but once OpenAI knows it gains a benefit by ordering items, how difficult is it to lift that restriction? Nobody knows. Might be easy, might be hard.
Some of those are very major, though, and I think the word "restriction" is a bit misleading. Having 5 invulnerable couriers is not really a "restriction" in that it limits or simplifies parts of the game -- it's just a fundamentally different mechanic that changes the way the game can be played.
> DotA players are probably the worst to ask, because the AI is already stronger than they are at a subset of the game and they have no insight into how quickly that subset could be generalised into other aspects.
I think that's a little unfair. Most folks have been pointing out that OpenAI's current momentum-based "deathball" strategy seems to fall apart without infinite regen and a limited hero pool, both facilitated by the current set of restrictions.
I'd agree that nobody really knows how well OpenAI will adapt to the full game, but I disagree that the criticisms I've seen are meritless. OpenAI's current level of play is definitely impressive, but I think there's still room for skepticism given the current restrictions. I (and I think a lot of others) would be pretty disappointed if the TI showmatch happened with the turbo mode couriers still enabled.
People said similar things about Go AIs and ko fights. And in the end it turned out that neural networks handled kos fine but ladders were a challenge.
On the deathball strategy in particular, consider that we expect a superhuman DotA AI to change the DotA metagame, so playing off-meta doesn't tell us anything. AlphaGo would invade 3-3 point a lot more enthusiastically than a human player. This was considered a classic beginner mistake for many years; now the theory has been readjusted to cope with the fact that AlphaGo stuck with it and just considered it a good move.
We can safely say that the courier change has made a deathball strategy more powerful and it seems quite likely it is not an optimum strategy. But we can't be sure until OpenAI tests it, and we absolutely can't be sure that OpenAI won't just learn a new style when the conditions change.
The criticisms have merit, but nobody has enough data predict anything about the future. Particularly a professional DotA player.
Totally agree. This one change alone is _so central_ to both the bots laning strategies and meta-game team strategies. They can't just leave heroes in lanes forever no matter what, and have all 5 heroes literally never go back to the well if there aren't 5 couriers. Not to mention their initial item builds, stats-only-4-man-the-lane-for-first-blood bullshit wouldn't work at all without constant ferrying of regen on the couriers.
>4400 MMR
OP is being extremely satirical here by the way. He means he's not great but knows how to play (and definitely not semi-pro) but that context might be lost if you don't play Dota!
>Is it more accurate to think of the bots as 5 separate minds, or a single mind controlling 5 heroes?
They answered this on the last stream, iirc it's 5 identical clones with the same goals, but not sharing any knowledge, info, or decisions with each other.
I'm just 3.5k (I think that's 70th percentile), but I know lots of 4.5k players. To describe the average 4.5k player: Probably has regular groups of people they play with at different skill levels (anywhere from 2.5-5k+), regularly plays battle-cup on Saturdays, maybe played amateur JoinDota league, maybe had a laugh and played open qualifiers only to lose in the first couple rounds, and probably log between between 10-20 hours per week into the game.
4.5k players know how to play to a very good standard and beat the vast majority of other players, but are miles away from the weakest of the professional scene. 4.5k doesn't even appear on the leader boards.
Solo - the guy that is the captain of Virtus Pro, was at 4k for the longest time. There's more to dota than just mmr.
There are players at 5.5-6k range that still do not understand the basics of team play, but are just extremely mechanically gifted and are in great gaming shape.
He also mentions that this was in the past. 4 or so years ago 4400 MMR was in the top 1% of dota players. MMR creep has happened significantly since then.
Go and look how it works: https://blog.openai.com/openai-five/
I asked about this in a previous thread[1] and received a response that a network blip caused all the players to drop from the game, in which case OpenAI Five was programmed to pause.
[1] https://news.ycombinator.com/item?id=17700001
EDIT: Fixed thread link.
They released their architecture: https://www.reddit.com/r/MachineLearning/comments/9533g8/n_o...
In the above thread, a reddit user noted that 512 out of the 2048 unit input into the LSTM of each bot is shared (max pooled across players).
This means they are telepathically linked and never need to worry about communicating, disagreeing, etc. They know how the others are interpreting their local inputs because it's explicitly shared. So it's not really fair to call it 5 separate minds.
You can't really call it a single mind either because if you assume the LSTM is the "mind" (since it's the only place that has memory of previous state), that state isn't read by any other bots.
This was done by the game coordinator -- the humans' machines DC'ed during the game. https://news.ycombinator.com/item?id=17700233
This is exactly what is happening with Go right now. Many pros are emulating and learning from AlphaGo (Zero) and are starting to play moves that were always thought to be suboptimal until now.
I thought they'd at least remove more of the rules (5 couriers, no illusions) or add some heroes.
Does the OpenAI team think there's a way to adapt the UX of DOTA 2 "Perfect Information Edition" to communicate the game better to human players?
Honestly I think that it would not be much more difficult to train a bot that looks at screen pixels and outputs keyboard and mouse events instead of using the bot API. In fact it might be easier to code, but the problem is it would require several orders of magnitude more processing power to train, which is impractical. I am confident it would work if the processing power was available, given the success of these techniques on other problems.
If they have to use the standard UI they lose a significant amount of information as they only view a very small percentage of the map per frame + the minimap. Just think about the implications for team fights. If all the agents have different information about a fight, then you aren't necessarily going to see the uncanny stun stacking and perfect long range nukes. The AI cannot assume that all the other agents share the same information, and so the other agents might not be as predictable. I imagine fragmented information would dramatically change behavior, for the worse. They will probably act more like human players - more cautiously and with more mistakes due to "miscommunication".
>"OpenAI Five does not contain an explicit communication channel between the heroes’ neural networks. Teamwork is controlled by a hyperparameter we dubbed “team spirit”. Team spirit ranges from 0 to 1, putting a weight on how much each of OpenAI Five’s heroes should care about its individual reward function versus the average of the team’s reward functions. We anneal its value from 0 to 1 over training."
The first is that the bot strategy currently revolves around the special rule of 5 invulnerable couriers. Bots find microing lots of units effortless, so the map constantly showed each bot's courier flying back and forth carrying regen. The bots never had to really go back to base or their shrines to heal. This is important because it changes the meta of the game entirely. The way the game is structured allows only one (very vulnerable) courier per team. Usually this means that after a team fight, teams need to reset since they've expended significant resources for the fight. But that meta was non existent under the rules for matches against the Open AI five. The humans had trouble coping with this as they weren't used to the idea of ferrying regen constantly.
Takeaways here - I could go on about the nuances of a single courier. But basically, the bots' gameplay will likely have to change once it comes down to 1 shared courier per team. Not sure how that will affect the architecture of a "no shared mind". Also, humans will likely need to take a page out of this gameplay and realise that couriers are a highly underutilized resource. Every second it's not doing something for no reason is just as bad as a hero not doing anything.
The second observation comes from the last game of AI vs pro humans. This was an interesting game where the audience picked a losing set of heroes for the team. Despite a predicted chance of winning being less then 2% (iirc) he AI could have probably won on account of being mechanically better than the humans. But their insistence on sticking to a strategy of "push hard" found them doing really strange things. The strangest of this was Slark running ahead to cut down creep waves in the lane on its own. The human players knew this would happen and they kept forcing the Slark to go hide in the trees and at some point they were always able to corner it and get the kill. Over and over again. The Slark never changed.
Similar things happened around the map during this game.
What should have happened was that the AI should have adapted to its disadvantage, and poured its efforts into first defending and then snowballing later with its mechanical advantage. But that element of "intelligence" was never there.
The takeaway is this. The AI will eventually beat the humans on account of them being always mechanically better. They need very slight changes in their strategy to win 99.9% of the time. They can be aggressive beyond any human possibility because they can calculate everything to perfection. How long it will take them to travel across the map vs how much longer it will take for an opposing hero to have its ultimate ready for example. There are a lot of mechanical components to Dota that the AI will always have an advantage over. But the AI will likely always reveal quirks that can be turned into dumb winning strategies (aka cheese strats). Something like the whole team fighting from the trees for example might just confuse the AI terribly. We don't know but every now and then someone will discover it and the teams working on the AI will have to "patch" the behaviour.
Final takeaway from all of that - I'm not sure if training the AI towards "objectives" is really the best metric towards making an intelligent bot. It seems like what's instead happening is that we get software that has no intelligence at adapting in the moment to things its never seen even if they are brain dead. But it'll get better at hiding them through mechanical perfections.
Upside - We get AI's capable of doing increasingly complex things in a seemingly perfect manner.
Downside - We get a scary future of AI filled with byzantine issues that need to be "patched".