How we built the unbeatable Dota 2 bot
blog.openai.com
blog.openai.com
Not to diminish the teams' achievement. This was awesome. But again, it's just more over-hyping of AI that should be called out for what it really is.
But that is not unfair. Humans receive plenty of curriculum training as well, we're not supposed to figure out the world by bumping into walls. Even in Dota2, the top players learned from observing each other how to deal with the bot. In fact, efficient retraining to include new strategies on the spot would be a very human-like learning ability.
I'm not a Dota2 player, but like SC2 for example is a game with LOTS of room for AI improvements. I've always thought that having some sort of APM limit might actually encourage AI authors to adopt new and unique approaches to macro-strats, but it doesn't seem to be on the horizon.
When it comes to do a small thing rapidly, I think bots are almost always going to win.
When it comes to do something large-scale with finesse, I think humans are going to have an advantage for a LONG time.
I think that part of what makes human agents so effective at certain tasks, especially in the context of being up against another human is that we can evaluate an event and better understand the WHY of it relative to the player that played it.
If I see a player pull back a bit, sometimes I think to myself that maybe they saw something they weren't expecting or something they weren't quite sure how to handle. When a computer sees the same move, a floating point number among millions changes slightly. I can try and figure out why they might be pulling back, if I did something weird or if I did something totally normal I might suspect it is bait, etc. I can think all these things in a short period of time and while large AIs might have better FLOPS than me, it doesn't understand what I'm doing, why I do it, etc.
Curriculum learning isn't as effective in bots as it is in humans is my contention, I guess.
Fair/unfair is a pointless observation when it comes to humans vs bots. The diversity of human-based problem solving is the perfect friction to train AIs against, imo.
What about this tweet from OpenAI?
"Our Dota 2 AI is undefeated against the world's best solo players" [1]
Also Musk called it more complex than Go,
"OpenAI first ever to defeat world's best players in competitive eSports. Vastly more complex than traditional board games like chess & Go." [2]
[1] https://twitter.com/OpenAI/status/896157788908290048
[2] https://twitter.com/elonmusk/status/896163163581825025?lang=...
I think there is a fair argument for Dota being vastly more complex than Go but there almost certainly isn't for 1v1 SF mid.
MC sampling is an essential tool (when used together with simulation and neural nets). Don't dismiss it yet, it has a lot to offer to deep learning. For example, if you don't have enough data, you can create it by brute forcing / simulation.
> Otherwise please use the original title, unless it is misleading or linkbait.
> Please don't do things to make titles stand out ...
(the original title is the much more sedate "More on DOTA 2"; in case it is changed by moderators, at the time of this comment it says "How We Built the Unbeatable DOTA 2 bot", which is not even a quotation from the article)
The submitted title has certainly shaped the commentary (see people reacting to "unbeatable", and whether it's an impressive achievement, vs reacting to the technical content).
So, in a nutshell, I think we should avoid titles like this, and I believe the guidelines transparently say so as well.
So whilst the new edited title is a bit click-baity, it is nevertheless more meaninful than the original.
It's also a lie, given that the bot was beaten multiple times by plenty of different people.
At the same time most of the deep learning field has been focusing on implementing super-human perceptual abilities (vision, hearing, translation) that come instinctually to humans. The higher level reasoning and memory/attention augmented machine learning is still cutting edge research. I think DeepMind and OpenAI are driving research towards that end.
I am not saying the bot should be able to play 5v5, I am saying that he ist trained to play this one very specific 1v1 scenario. He is doing really well with that, but as soon as you pick a different hero, the bot will struggle.
My favorite part:
The first step in the project was figuring out how to run Dota 2 in the cloud on a physical GPU. The game gave an obscure error message on GPU cloud instances. But when starting it on Greg’s personal GPU desktop (which is the desktop brought onstage during the show), we noticed that Dota booted when the monitor was plugged in, but gave the same error message when unplugged. So we configured our cloud GPU instances to pretend there was a physical monitor attached.
Dota didn’t support custom dedicated servers at the time, meaning that running scalably and without a GPU was possible only with very slow software rendering. We then created a shim to stub out most OpenGL calls, except the ones needed to boot.
So many things don't fit in.. but that might be because I am neither good at DOTA 2 nor at AI stuff. Good luck with the 5v5's openai. And if the AI bots become public, maybe i'll get to play against them some day.
If you tell the API to make you attack a thing once, you're 'attacking' until the projectile is created, and then you're 'idle' immediately afterward.
So long as you're feeding the hero actions, it'll always animation cancel.
What I think happened here: the bot lost against someone using an item that hadn't been whitelisted before. The devs decided to include the item among the possible choices for self-training. Exploring the new choice, the bot adapted to using it to defeat its previous strategy. As a result, it now had to change its strategy to account for an opponent using the item.
So it's kind of interesting and equally scary that we are looking at a potential future, possibly near, where most of the work can be handled by a trained robot with humans only supervising them in exceptional cases. Robots can then could be only need to be smart enough to alert the human supervisors when they reach a state where they cannot perform the job normally, for e.g. power grid going off in a factory and they couldn't figure out what to do, it could only be a 0.001 % chance. The point is AI need not get as good as humans in learning complex scenarios to replace majority of us, it just need to do 99% of what we repeatedly do, much better than us.
How does the bot understand the value of long-term strategy?