OpenAI’s Dota 2 defeat is still a win for artificial intelligence
theverge.com
theverge.com
The bots excelled at things you expected computers to be good at, consistently quick reactions/mechanical skill and coordinated skirmishes. The bots fell flat on strategy, optimal resource usage (short and long term), and most importantly, coordinated decision-making.
So as an result, they dominated the early game which has always been heavily mechanics favored and absolutely fell apart mid/late-game when decision-making mattered (and often made decisions that benefited early game economy at the cost of mid/late game).
Computers being good at mechanics is not an impressive result.
With all due respect, the professionals who have played against the bot have been pretty impressed with previous iteration. The major roadblocks at TI were probably that the bot did not efficiently learn courier mechanics and that the draft was not done by the bot.
A human has ~200ms simple reaction time, then begins to move their fingers. Which then moves a mouse to target an area, which then requires hotkeys to be pushed. etc. etc.
Its still definitely a mechanical advantage, and it seems evident in the gameplay videos. The bots still seem "super-human" in their aim and some of the timing of spells.
What was interesting to me was that since the bots trained against themselves, they discarded this combo as being ineffective - the result being that they did not use this combo even against vulnerable human players.
This means that Dota 2 was balanced around the expectation that human's best reaction would be at least 400ms, because I've never seen Axe's blink calls being dodged unless the player had vision of Axe before the blink and thus was expecting it.
Multiple invincible couriers makes the game very similar to the 1v1 show match last year where the strategy of just out-live the opponent wins. The casters have never played that version of the game before, and wouldn't know to use that strategy.
The more interesting part was, the bots attempted to use that very same strategy in the single-vulnerable normal version of the game which was drastically less effective; you could only ferry out consumables at 1/5th the rate as before and the courier could be killed so it wasn't a riskless strategy like it was before.
I 100% disagree that the invincible couriers make the game "very similar" to the 1v1. There's many more layers of engineering that went into that version. The bots had to learn drafting, cooperation, highground pushing, Roshan, vision, invisibility, etc.
I'm amazed you can trivialize all this as a 'joke'.
In my opinion, what i see is a very good player who knows how to chain stun precisely without any strategic depth. If you claim to have built an AI system, which you ultimately want it to evolve to AGI, you at least expect some sort of strategic decision making at the macro level. Though since it has almost perfect micro, it can easily outweight the most of teams. So yeah, with that expectation I see this as a joke, too.
P.S. The model is trained with 128k cpus and 256 gpu. It is able to play 180 years worth of game in a day. Think about it.
It's the first line of the article: https://blog.openai.com/openai-five/
>Our team of five neural networks,
They use a hyperparameter called team spirit to cooperate. I don't think the goal of this is AGI at all, so I don't see why people are making that leap. But sure, for the geniuses of HN this must clearly be trivial.
I think they should have trained with 1 courier right from the beginning.
If I had to guess, they started out with the 5-courier setup because otherwise the bots would end up fighting over the courier -- just like a typical pub team. :)
The way I would guess this played out was that OpenAI tried multiple variants of the game, and this version ended up being good enough to beat human players. Of course, they later learnt that the bots were winning not because they were good, but the variant was imbalanced and its strategies unfamiliar to humans.
Deathball strategy has been a thing since even before 2014...
You have argued yourself into a corner. It happens.
No. We had implemented scripted courier logic for 1v1, and when switching to 5v5 the easiest starting point was to run five copies of that logic. (Dota's Turbo mode also has five invulnerable couriers.) In June, we had more important restrictions to remove — such as wards & Roshan; you personally were focused on the particular heroes we'd chosen: https://news.ycombinator.com/item?id=17392455. The couriers only become our most important restriction in August.
Of the list of skills you mentioned that the bots "learned", on the main stage, the draft was provided by humans, and from a significantly limited set, the bots never pressured highground, checked for Roshan when it was impossible for it to be up, had nonsensical placement of vision (placed vision where vision already existed), continuously invested in anti-invisibility consumables even when the opposition team had no invisibility.
If the mechanical skills were toned down to human levels, I don't think the latest iteration of the bot on the latest ruleset can even compete with the average human player, much less the pros.
The games they won weren't actual dota games - they were a simplified subset of dota that overweights teamfights and mechanical skill - exactly the subset that we'd expect an AI to be best at.
It's argued that drafting is one of the major components of competitive play and more often than not, draft advantage wins games, but that wasn't a factor in this match.
Only at human speeds.
Once you break the 10,000 APM barrier however, Zerglings start to dodge-tanks and other shenanigans begin to happen.
https://www.youtube.com/watch?v=IKVFZ28ybQs
Micro has a general advantage of growing at a rate of N^2. That is, 10 Zealots, with perfect micro, can perform roughly as well as 100 Zealots with the worst micro possible.
Ex: 100 Zealots come in one-at-a-time vs the 10-zealots in an inverted-V shape. The 10-zealots will defeat roughly 100.
The bigger and more complicated the board gets, the more and more micro becomes favored.
This article also really understates how clowny the bots were playing. They appeared confused and aimless for much of the game, until they found a clear objective like an enemy hero that moved too close to them, or they had a numbers advantage that allowed them to pressure an enemy building. They were brutal at closing gaps and pursuing enemies to kill without hesitation, but didn't perform well at all beyond that.
This isn't to take away from the accomplishments of OpenAI - to get to this stage after 18 months of work is an impressive engineering feat. It's just telling that the bots have only learned one single strategy, don't even execute that very well, and still have to play with hero restrictions that rule out counters to that specific strategy of aggressive skirmishing and pushing.
I kind of liken it to a human who learns math - you start off by learning the rules (addition, subtraction, algebra, etc.) and work within those confines before you start to use these basic rules to push the boundaries and explore (proofs).
I suspect this has to with the fact that the bots were designed to be agents with no leadership. As you say, the bots performed very well in skirmishes, they lost because they didn't really know what to do beyond that.
They might fare better if one of the bots could act as a team captain and provide high level instructions to the team, i.e., assign people to harvest, defend, or push.
It seems like that's what the human teams have over the bots -- somebody telling the individual team members what to do.
OpenAI 5 had a number of glaring flaws in what it learned. It seemed to have never come to the conclusion that certain characters would have better outcomes if they focused on certain activities. Every character farmed, any character would grab the Aegis (reward item) after killing Roshan (usually this is reserved for a character likely to be in the middle of the fray, dealing a lot of damage).
Another problem was that, until the last few weeks before its matches against pro players, they had allowed both teams to have 5 couriers. This allowed OpenAI 5 to keep up relentless pressure by constantly bringing themselves consumable health regeneration items. This was so unlike a regular game of Dota that the community (including their opponents) complained because normally the courier's time is a valuable resource, and the courier can normally be killed if it is used too close to enemies. With an endless supply of healing items with no risk involved, it didn't even resemble a game of Dota. They did away with this for TI, and it revealed a serious weakness in their relentless-aggression strategy.
I get the impression most of OpenAI's games vs itself ended very quickly, with one team making a mistake during the relentless-aggression early game pushes. As a result, it seemed to have no idea what to do differently as the game went on.
I'll own up to wanting to chalk this up to "AI wisdom." The feed-the-carry strategy has always seemed like precisely the same kind of premature optimization that makes AI often easy to beat.
> First, OpenAI could have created a vision system to read the pixels and retrieve the same information that the bot API provides. (The main reason it didn’t is that it would have been incredibly resource-intensive.)
Mmmm, no. Reading the pixels would not retrieve the same information that the bot API provides, because the screen does not contain all that information at any given time.
Any Dota 2 player will know that they don't have access to all this data simultaneously:
https://s3-us-west-2.amazonaws.com/openai-assets/dota_benchm...
More complicated is the logic to know what to look at when, by keeping track of which parts of the state might be important and out of date, and controlling the viewpoint to update it. I haven't seen a good example of this.
I agree that learning how to retrieve the info would be another rather tough problem (on top of all the tough problems the current incarnation of OpenAI's bots already need to solve).
https://github.com/cshenton/atari-leaderboard
And that list is incomplete if you want to include human high scores on non-emulators (not included in that list). This is even after the reinforcement learning algorithms have been given orders of magnitude more training time than humans. Furthermore, much of the machine performance over humans can be attributed to better reaction time.
IIRC, you can easily build a super-human Atari bot. Just add some off line planning (e.g. paper by Guo), or add manual rewards or features.
No one is doing since there is no point. We want to study Atari WITHOUT these "cheats", so that we can then apply these algos in more complex situations.
For generalized reinforcement learning algorithms. Humans almost always beat the "machine". The only cases where machine wins is on 100% hand eye coordination and duration.
All ML algorithms are terrible at figuring out how to plan de-novo. So anything that requires multiple contexts or planning is a fail.
Yet on the other hand, it is apparent that reinforcement learning algorithms have surpassed humans in board games such as Go and Chess.
I saw a strange behaviour. In the second game, I saw witch doctor use "Maledict" on neutral creeps, but this skill only affects enemy heroes. This only waste mana, put the skill in cool down and have no benefit. How can AI learn it?
As a further analogy I often see humans make ridiculous and insane leaps of logic (how assignment works when you first learn programming; free speech applies to things besides governments; etc.), and we run on the same style of hardware.
To an extent this demonstrates how far we are from generalized AI. All the Open AI system is is a very specialized animal, that's lived through millions of simulated generations very quickly. It doesn't understand language or concepts - we can't even begin to tell it the rules of the game, only though simulated evolution can we teach it - or even the self awareness many animals do. It's likely closer to the nervous system of an insect than it is to the brain of a dog.
Could we please lay this empty platitude to reast? "Planes don't fly by flapping their wings like birds, so why should computers think like humans?".
Well, except that we were only able to built flying machines (rather than floating ones) when we figured out why birds can fly [1].
And then, the whole of our computer science is based on the idea that a computer is, actually, the same kind of system as a human mind- a computational device, a machine that can compute everything and anything that can be computed. This is the deep insight that informs the ambition to create artificial intelligence: that brains are a kind of computer, computers are a kind of brain, and they can both compute the same kind of program, in other words- intelligence.
Though we may not know how human minds work exactly and therefore we can't readily copy them, it is thanks to Turing's and Church's insight that we even have computers today. And so, in a very real sense, we can only ever create thinking machines that do intelligence in the same way that humans do intelligence [2].
_______________________
[1] I mean, a boomerang is really an airfoil so I guess Australian Aborigines had figured out flight long before the Wright brothers, but I gues they didn't need flying vehicles back then?
[2] Unless there is a paradigm shift and we figure out a better way to do it etc etc disclaimer disclaimer.
Formal logic and the Turing machine were the model for computers.
Turing and Church did not invent computers based on their careful study of human biology and physiology.
Obviously, the Turing-Church thesis didn't have anything to do with brains. It does, however, have very important implications about our ambition to create human-like artificial minds.
Also, I'm sure that theoretically optimal "intelligence program" on silicon processors is vastly different from organic ones. They may share a lot, but the constraints are just so vastly different including the most important ones like energy efficiency or space availability.
That's not impossible! For example, it's easy to find algorithms that a computer can carry out without error that a human mind would really struggle with. Although this is a case of computational resources, rather than the expressive power of the computational apparatus, it's still the case that it's not always possible to find programs that both humans and computers can compute in practice.
So it may even be that, while computers can run some programs very efficiently, that humans can't, it's the other way around also and computers can't efficiently run the programs that humans can.
In which case of course, either the entire AI enterprise is doomed to failure, or we get lucky and there is some other way to do intelligence that is available to computers but not humans. Who knows!
Our current computers function almost nothing like human brains.
All we can say, is that we can write some programs now that give similar output to humans for certain well defined, limited problems. But even then, the mechanism via which those outputs are produced are very, very different.
The mistake is assuming machines that can perform tasks better than humans, will be essentially "like" humans in some fundamental, profound way. There is no reason to believe this is true. These machines may be able to mimic us very well, but their true underlying nature, their goals, ambitions, subjective experience, and values could be very, very different from us in ways that could be very unpleasant for us.
With respect, but I never said anything about computers functioning like human brains. What I said is that computers and minds can compute the same class of programs. I make no assumptions about underlying function.
>> The mistake is assuming machines that can perform tasks better than humans, will be essentially "like" humans in some fundamental, profound way.
I never stated any similar assumption either.
I think you're responding to some opinion that I didn't express.
> [1] I mean, a boomerang is really an airfoil so I guess Australian Aborigines had figured out flight long before the Wright brothers, but I guess they didn't need flying vehicles back then?
What? I'm not sure anyone would count a boomerang as "figuring out" flight let alone how birds fly.
I'm no historian but I'm pretty sure flight was first figured out mechanically through trial and error well before we had a good theoretical model for flight, let alone figuring out how birds fly. Either that or we have drastically different meaning for what it means to figure something out.
It took ages for us to fully understand how flight worked, well after we could fly. That's the basis of the apocryphal story:
https://www.snopes.com/fact-check/bumblebees-cant-fly/ https://rationalwiki.org/wiki/Bumblebee_argument
Fair or not fair, there is no API to the real world. If we are going to create machines that can learn and think outside of simulations, however complex they may be, these machines will need a way to interact with the physical environment.
Gary Marcus is on the right track here. Building systems that can "handle the complexity of the real world", as per OpenAI's stated goal according to the article, is incredibly ambitious and pointing out the vast distance separating the current state-of-the-art from that lofty goal is, well, only fair.
I mean, if you think about it, back in the '70s, in the original AI Winter, one of the big criticisms of AI research was that it languished in simulated environments like blocks world and didn't perform nearly as well in the real world. And here we are today, celebrating a bright step on the path to conquering yet another simulation.
As a result, training in simulated environments doesn't help handle the full complexity of the real world. Even if we had robot bodies that could move as freely and manipulate objects with as much dexterity as they can do in simulations.
But you are right that it will fail on anything that is hard to simulate. And a lot of trivial things are very hard to simulate.
Tesla recently failed at things like plugging two cables together. Picking stuff like bags at Amazon warehouses may be another example.
> At least one previously undiscovered game mechanic, which allows players to recharge a certain weapon quickly by staying out of range of the enemy, has been discovered by the bots and passed on to humans
I assume they're talking about blink dagger, and something more advanced than "don't take damage"?
it was discovered that if you stay out of vision and cast raze, the other player does not get stick charges. that was the 1v1 shadow fiend bot a year ago though.
"Sumail pointed out that the bot had learned to cast razes out of the enemy’s vision. This was due to a mechanic we hadn’t known about: abilities cast outside of the enemy’s vision prevent the enemy from gaining a wand charge."
That's not an undiscovered mechanic in the world of Dota - that's been known for a while and at least documented since 2015 https://dota2.gamepedia.com/index.php?title=Magic_Stick&oldi....
I wouldn't be surprised if it was known before then, I certainly remember this from a while back.
Again, if that's not what it was, then I take back what I said, but if it was, I do think that statement is misleading as written.
Oh, yes, indeed, there are limitations. The robot hand in question can only manipulate cubes and then only a specific kind of cube with standard dimensions, as far as it's possible to tell from all the demonstrations publicised by OpenAI. And, I'm guessing, if they had a robot hand able to play the yo-yo, they wouldn't hesitate to show it.
Don't think match win would really mean something. Only from the marketing perspective.
The announcers/casters were a bit ambiguous about this, AFAIK they said something like : "the teams agreed on a predefined, balanced list of heroes before the game...", and there was no drafting.
(I'm not native speaker, and I don't know much about Dota2 rules)
AI has long ways to go before it can defeat humans in a complex game like Dota2.
https://www.youtube.com/watch?v=IKVFZ28ybQs
That's the "gloves come off" AI. Granted, its a problem solved by raw, mechanical button-pushes at a rate beyond what is possible on a keyboard/mouse, but it demonstrates the greatest skills an AI or computer has.
Of course, no one wants to see an AI beat a human using known methodologies. They want to see an AI win at the abstract game of "a game of strategy".
From this perspective, I think the poker AI results over the past year or two have been far more interesting. AIs figuring out the Nash Equalibrium and playing bluffing games around it.
Sure, its a bit mechanical and mathematically inclined, but generalizing and solving the bluffing game is IMO far more useful.