OpenAI bots competing against Humans right now
twitch.tv
twitch.tv
I wonder if this is an artifact of the training methodology: maybe if your team is very weak then your choices are also weaker, and reinforcement learning doesn't work as well?
When the win percentage for Go AIs gets to around 5%, every action it can take results in a losing game so it can't make the difference between normal play and super strange moves anymore.
When every choice is really bad, humans tend to still go with their normal strategy and wait for their chance to turn things around, but bots assume the opponent is playing perfectly, so they act like their winrate is going to stay near zero no matter what they do.
Several levels of weak opponents should be used, with varying probabilities, to tune the AI’s robustness against real-world, imperfect competitors.
Taking a tier 2 tower nets everyone on the team 120 gold (a further 150+ gold goes to the hero who gets the last hit), and losing Sven probably gave the opposing team less than was gained.
Perhaps the AI simply placed more value on increasing the total net worth of the team than it valued saving the life of one of its core heroes. Additionally, there was no guarantee that he would have been able to escape, as Sven was deep on the enemy's side of the map, and there's a very real possibility that he could have been ganked from someone in the jungle had he attempted to retreat.
- Potential chance the enemy team would deny the tower before another friendly hero could take it (netting Sven's team 0 gold for the time spent whacking away at it)
- Map vision (removing a T2 often cuts a significant section of map awareness away, since the tower is no longer providing vision or protection)
- XP gains (Sven won't gain any XP while dead, nor from killing the tower)
- Creep equilibrium (this is less important, or at least thought about less often, later on in the game and past T1 towers, but might've been a factor in drawing the creep clash point to a particular location)
- Dictating team net worth averages (to some extent, if they predicted a loss in opponents forcing a teamfight or predicted a likely pickoff, gold lost could be minimized now by taking a death early, lowering the average net worth on the team).
Obviously, there are others and these can also be mixed and matched in various ways (e.g. cutting off map vision so they can more safely farm additional jungle creeps).
Not saying any of these aspects _were_ a part of the decision to trade Sven for a tower, but.. just wanted to include a few more subsurface aspects that _could_ be used in such a decision.
I'm not sure if the AI can surrender (I only managed to watch the first two games as it was rather late at night) but it might be a path to explore; having the AI give up if the game cannot be won anymore.
I wonder if this situation can be fixed by adding more randomness. For example, force AI'1 to be in a losing position to AI'2, but then suddenly switch the power level of AI'2 to be much weaker (where mistakes happen) so that AI'1 learns how to fight its way out of tough situations.
These are the kind of actions you specifically don't want to code in because you're throwing in human knowledge. You want the AI to learn by itself that using anti-invis when everyone is visible is a low-value move.
The purist in me was even mad that they had a hand-crafted evaluation function. (e.g. prefer gold, prefer taking towers, each given some arbitrary value)
"We’ve increased the reaction time of OpenAI Five from 80ms to 200ms. This reaction time is much closer to human level, though we haven’t seen evidence of changes in gameplay as OpenAI Five’s strength comes more from teamwork and coordination than reflexes."
Dota & HON had people mod their client to give an optional bigger FOV resulting in bans for cheating.
I'd assume the bots don't have to specify their screen position, plus no orientation response means a limitation on this wouldn't be meaningful anyway. What I'm saying is there's a big difference between 'I see lion on the mini-map down there' and 'lion showed for 1 frame on the other side of the map, his HP is 324, he has a TP scroll, no boots and a health pot.'
Something I noticed is the bots seem to like range and AoE far more than the normal human meta. The humans being limited in the distance they can see to one screen were frequently just failing to appreciate how dangerous 2-3 bots half a screen away were to them.
Quite a few teamfight wins came from the bots inevitably causing far more damage to the entire enemy team via heros like DP & Gyro. But this isn't really perfect teamfight execution. I'd have really liked to see a mirror match.
Well, maybe not technically, but it does make it much much easier to take in all of that info and process it. You can't expect a human player to keep a perfect record of all heroes' hp, mana, all damage being dealt, all abilities being used etc. during a chaotic fight, yet the API yields this information effortlessly.
> The micro- and reaction time edge is also being dulled to being more human-like.
And yet, the bots showed superhuman near instant reaction times. 200ms is a very low amount to process very complex/confusing audiovisual data and react precisely.
Regardless of if the information comes from the machine viewing the damage count and knowing exactly how much HP a given hero has at that level/gear/just by looking at the bar, or if the information comes from an API, the machine has a perfect memory of this and all other variables, whereas humans don't.
Humans, not so much, as in all top-level competitions, human abilities improve minimally at the top, because we have millions of humans competing against each other until the plateau of human performance is reached. Then you can push that a bit more with drugs (see doping in sports). And after that, you are pretty much done.
So it's only a matter of time and effort until AIs are fully unbeatable.
That's just not true in doto. The player base improves quite a bit over time. The top pro plays from only a few years ago are not impressive anymore.
And if watching the screen, do you want it to have bad eyes like we do too (good resolution only in the center)?
I mean, "human vs AI" matchups are ostensibly about strategy - machines already win at timing and twitch, there's nothing to test. But esports games aren't pure strategy, they all involve various amounts of timing, twitch skills, the ability to monitor lots of details at once, etc. Those are all things that AI opponents can (trivially) do perfectly, which gives the AI a huge advantage. It then follows that an AI player should be able to win even with an inferior strategy (which makes you wonder if these games are really suited to AI research in the first place?).
[1] For an entertaining case-study, check out Day[9]'s learns DotA2 series.
It occurs to me that for a really even playing field, the humans should probably be allowed to make and install UI mods if they want to. E.g. if there's an advantage to using an ability precisely when your hit points hit 50% (or whatever), an AI can easily do that reliably so the human should probably be able to if they want to.
(Of course, for heavily twitch games like Counterstrike, being allowed to use UI mods (i.e. aimbots) would break things. But then, I suppose that the extent to which UI mods break a game is more or less the extent to which that game favors twitch over strategy.)
Yes, it is. Also, “seeing” the screen rather than being able to directly introspect the game world digitally. Orders of magnitude harder. This is known as Moravec’s Paradox.
Like, imagine if this was a chess AI, and we were trying to determine who was better at chess, humans or AI. Would you make the AI use robotic hands to move the pieces? No, because thats not the interesting part of chess. The interesting part of chess is the strategy.
And consider a game like Quake where mechanical skill is even more important (even though mechanical skill matters, Dota 2 is still primarily a game about strategy and team coordination).
Another interesting part will be creating an AI / neural network that can utilize inputs that are closer to human level inputs (e.g., using the frame buffer and audio out as input to the neural network and passing the outputs of the neural network to a keyboard and mouse driver). Just let the network train itself without having a human laboriously determine the topology of the neural network. Such a neural network can then be applied to several different types of games / problems much more quickly than at present where significant human labor is required to generate deeply customized neural networks for each game / problem.
I agree it'll be even cooler when it all justworkstm end to end, but in terms of incremental 'holyshiticantbelievethatworked' this is at least as big a step as it will be when they add in direct visual input.
One of the next significant moments could be taking the current Dota 2 algorithm and massaging it to use human style inputs and outputs. Please correct if needed, but the current Dota 2 algorithm boils down to (1) a fully connected network that generates an input state vector from the Dota 2 bot output interface, (2) an LSTM of sufficient length that generates an output state vector from the input state vector, and (3) another fully connected network that generates the Dota 2 bot interface inputs from the output state vector. This could be updated to have (1a) a convolutional network that feeds into a fully connected network, where the input to the convolutional network is the frame buffer (and perhaps the audio output) and the output of the fully connected network is the input state vector, (2) the same or similar LSTM network, and (3a) a fully connected network that outputs keyboard and mouse commands instead of DotA 2 bot interface inputs.
It is an open question as to whether current compute power is sufficient for this massage.
I hope this little koan illustrates that this sentence is impossible to execute. The human always has to specify something.
--
In the days when Sussman was a novice, Minsky once came to him as he sat hacking at the PDP-6.
"What are you doing?", asked Minsky.
"I am training a randomly wired neural net to play Tic-tac-toe", Sussman replied.
"Why is the net wired randomly?", asked Minsky.
"I do not want it to have any preconceptions of how to play", Sussman said.
Minsky then shut his eyes.
"Why do you close your eyes?" Sussman asked his teacher.
"So that the room will be empty."
At that moment, Sussman was enlightened.
The exercise then becomes one of finding the minimal constraints needed to achieve the desired results. Please correct if needed, but looking at the Dota 2 neural network [1], it boils down to generating an input State vector from the Dota 2 bot output interface, running the state Vector through an lstm (of sufficient length) to generate an output State vector, and generating the inputs for the Dota 2 bot input interface from the output State vector. Update this network (1) to have the input State Vector generated from a convolutional network that feeds a fully connected Network and uses the frame buffer as input and (2) to have the final outputs of the neural network be keyboard and mouse commands instead of dota 2 bot input interface commands, then let the network train itself. The number of elements in the state vector, the number of convolutional layers, the number of lstm layers, and the number of layers and elements in each fully connected hidden layer could each also be determined by a recurrent neural network.
[1] https://towardsdatascience.com/the-science-behind-openai-fiv... (see the image under "The Architecture")
[ random capitalization powered by Google speech dictation ]
What if we could pit humans and ai bots at the speed of human imagination?
One could imagine a much better way of testing hand eye coordination, through a serious of mazes or puzzles or reaction tests.
It would be like trying to test hand key coordination by having a robot play physical chess against a person.
Am I missing something, or does that set consist of Checkers, Chess, and Go so far? (presumably with analogous misc games of comparable complexity)
Discounting the reaction time wins, I'd say the sample size is too limited to generalize to eventual AI behavior in more complex / open-ended games.
Extrapolation was the cause of the last AI winter.
So that's a pretty different game, it's got a big luck factor and has asymmetrical information and still the AI just kept getting better and the humans... didn't
I can see following hypotheses (in no particular order):
1. Human brain is the optimal solution in the space of all computational devices capable of playing games, and we can only approach it.
2. To do computation human brain employs some physical processes, we will not be able to replicate in the foreseeable future.
3. Human brain do not produce general intelligence, so we will not be able to replicate it as such a task is outside of the scope of our limited intelligence (while playing games isn't).
4. Human brain uses metaphysical abilities to do cognition, we will not be to replicate them at all.
5. Human brain is a local optimum, but the space of potential AIs' constructions is too huge to explore in the lifetime of our civilization, so we will be stuck at this local optimum with marginal improvements.
I don't see any of them as sufficiently likely, but your mileage may vary.
I believe there exists a combination of hardware and software capable of beating humans in all games. However, I also believe victory in a single game gives us minimal information on whether or not the system generalizes to many games (to say nothing of non-game, e.g. more complex, ruleless systems).
The reason I point out the unfathomable numeric complexity is that it makes the games, from the perspective of an AI, effectively infinite. AIs are calculating, but to an extremely superficial degree relative to the depth of the game. E.g. - when a chess program says it's calculated to 30 ply (15 moves for both sides) what it really says is that it's seen up to 15 moves deep after intentionally ignoring or pruning 99.9999999999% of moves which it thinks probably aren't good -- something it still often gets wrong, but its 'understanding' of what is 'not wrong' is strong enough that it still results in a phenomenally strong level of play, compared to humans. There's no doubt that perfect play in chess would still go 1 billion - 0 against something like AlphaZero.
So what matters is not the number of decisions to be made but the individual complexity of the decisions to be made. And in most games we consider complex the individual decisions are not really that complex, and complex systems can often be broken down into very simple games. For instance a great example of this is a 4x game. Taken as a whole they seem complex, but they're really just a large number of relatively simple components that are mostly independent. E.g. - Given this state, where do you explore next? Given this state, what do you research next? Etc. Another benefit for AIs in that in games we consider more complex, the value of any given mistake often becomes diminished. If you make a single bad move in chess, it's enough to lose the game. In a 4x game the weight of individual decisions is not so high, it's all about the big picture. But as perhaps computer success in Go shows most clearly, actually seeing the big picture is not really necessary to produce play like you do.
This, I think, is why research has moved more onto real time competitive games. Crushing humans at chess, go, and now poker as well is a pretty solid proof of concept for computers beating humans at any turn based game. When you start adding bunches of different layers to games I think it's more likely to handicap the human than the computer. Imagine playing some sort of 100x100 chess. We can only speculate, but I imagine the distance between the top AIs and humans would be far greater than it is in 8x8 chess.
I would disagree with this characterization. I believe at the time, it was (a) a problem that a machine had not yet conquered, (b) a problem that it seemed feasible that a machine might conquer, and (c) a problem that, once conquered, would point the way to general artificial intelligence.
I would point at (c) as the assumption that proved to be erroneous. Deep Blue was clever algorithmic and hardware engineering (with a healthy budget) but led to... what?
AlphaGo is a fundamentally different approach, which shows signs of being more adaptable.
Point being, that winning a game is not sufficient evidence that a given approach will scale to winning all games, much less generalized intelligence.
To put it in terms of the fallacy I read in an article linked on HN (paraphrased), 'The public assumes that if a machine can perform a task that humans can perform, the machine must be human-like, and therefore able to perform all tasks that humans can perform.'
But in the same way that we use rendering tricks to go beyond-state-of-hardware-art in graphics rendering (by abusing hidden limitations), so do we often build ml systems.
I believe the most optimistic point against me was the slide in this year's GTC keynote pointing to the "Cambrian explosion" in the diversity of ml approaches this time around.
There will be a time where we learn from the AI and the AI learns from us, where we trade victories and defeats as we adapt to each other.
Don't discount the ability of humans. They figured out how to exploit the 1v1 bot in a few days and soon humans had a 100% win rate using that strategy.
given more practice bots would beat humans. that's the point, train bots, which are faster to train than humans to beat humans.
To "beat" Atari games, AIs trained using reinforcement learning had to put in significantly more than the 10k hours one would expect a human to put in in order to become expert. So, AIs won't beat humans in tasks where training data is costly; however these cases are not interesting to researchers and hence you won't hear from the respective results.
Cases it already happened: board games such as Chess and Go, Poker, diagnostics of certain diseases using medical images
Cases where AI is still clearly inferior: video understanding, natural language understanding, motor control esp of hands and legs, general medicine, driving
Hard-to-classified cases (AI is better for some instances, worse for others): image tagging and classification, speech recognition (speech-to-text), diagnostics of certain other diseases using medical images (which might need to take into account other information outside of images)
More examples esp counter examples are welcome.
This is a misnomer.
The only variant of Poker where AI beats humans is heads-up (two-player) variant, which is simplest form of poker and also rarely played. The AI was (marginally) beating humans there by playing a game-theory-optimal strategy. For poker games with 3+ players, the GTO strategy (Nash equilibrium) no longer exists, so AIs need to use more standard techniques (search-based, reinforcement learning etc.), which are, at the current state of the art, laughably weak at poker.
Not to mention, that in poker the actual hierarchy of players' skill is not 100% obvious. You could distinguish at least two areas of skills:
- play vs other experts
- play vs amateurs/weaker players. Here' the goal is not to come out ahead (which, in long term, is a given), but to _maximize_ the dollar amount taken from these players, which is a skill in itself.
The observation applies to a given variant of poker (or any other domain). So if an AI beats humans with 10000-hour experience in that variant, the best experts in that specific variant are not far-off targets.
Superficially similar problems might in fact require very different techniques to solve as your example illustrates.
Edit: I looked up how much time it takes to train: "OpenAI Five plays 180 years worth of games against itself every day, learning via self-play." [1]
I believe there's a connotation with problems, that if the best algorithm to solve it is exponential (brute-force search), then we truly don't understand the (underlying structure of the) problem.
I have a feeling that it is not possible to reduce algorithmic complexity of finding optimal solutions in most of the intellectual tasks (those that are in NP complexity class and above).
Most likely it is a trade-off. Quickly cobble up suboptimal strategy / build better strategy from scratch avoiding all time-saving benefits of using known parts, and avoiding all the pitfalls of not reevaluating utility of those parts in the current situation.
AIs surely will need to use all the spectrum to compete with humans.
There's some of this in dota, but there's a cap on the skill level for most playable characters that pros generally get "close enough" to, and beyond that the strategic depth comes from area control decisionmaking. Theres over 100 heroes and many of them have really weird abilities, like the possibility of creating a temporary wall (earthshaker) or the ability to teleport anywhere on the map every 20 seconds (furion). I could be wrong though, maybe the AI is winning games by playing heroes with long range and perfectly microing them to harass and prevent the other team from ever getting gold/xp.
As somebody who plays StarCraft casually (gold/low plat in ladder), this is not true. It's even less true for pro players. The level of strategy in StarCraft is impressive, it's really hard to guess in which direction games will go when two very good players are playing against each other.
Sure, perfect execution when it comes to one strategy (say, mech-heavy Terran) will give you the largest advantage against your opponent, but failing to scout appropriately and guess what your opponent is up to means your strategy is dead. You also have to decide when to attack, how much you're willing to sacrifice to damage somebody's economy, when you want to focus one economy vs building units, ...
The video you sent with zerglings is a gimmick made for fun (it's a hard-coded AI using the siege tank's aim logic to divert zerglings from that). That would not win you a game. (because most likely a pro Terran would have destroyed your base before that)
In a way, the built-in AIs "macro perfectly", but they are terrible at strategy and fighting (because even fights are not just a matter of gimmicks, you need to split units in a special way, send diversions, attack at the same time from multiple fronts, etc.)
I do think a simulated reaction time or limited actions per second might be more fair, if they don’t have that already..
Like I could imagine OpenAI getting stuck in a subset of the draft pool for which it trained against, like maybe the top 10 of 18 champs. And then picking outside of that meta causes it to fall back on much less robust training/strategy.
When I think of AI, I think of something crawling its way out of purposefully adversarial situations such as this one. I would have loved to see optimal play from 5 wacky heroes.
I just have this suspicion that that wasn't optimal for that team comp.
But of course the matchup itself is a thing.
Ideally the pause would do nothing to the game. (It could be that the game has some glitches but let's assume it's solid.) Then one could assume that the pause does nothing regarding the AI players. But that is not true because the AI itself is not paused (I assume) and it has to handle this unusual input. This might throw off some versions of it and it probably learns how to handle the pause and how to use it to gain a small advantage. So it's all part of the game. Similar to when a human team uses the pause for tactical purposes.
- Item-usage for things such as smoke and wards (which were recently added to their reportoire) are not well captured by the bots yet. And the buying of wards were confirmed as a scripted event. It seems hard for them to capture the long term sparse reward of these. Smoke might not be needed by a perfect agent, but wards should be. The developer interview noted that it's not very clear what the reward for an agent warding even should be.
- Some of the big advantages OpenAI gains are in team fights, where it's 200ms reaction time (upped from 80ms to resemble human reaction time more) still strikes me as something that tilts it into a solid mechanical advantage. On several occasions OpenAI's Lion managed to disable a human player performing (what I assume to be) a move which shouldn't be interruptable. (blink-->shift+ctrl ultimate ability on earthshaker) This could have tilted teamfights in the human team's favor a few times if it wasn't stopped by the "machine-like" reflexes of OpenAI.
- Positioning before team fights by OpenAI are scary. There are very few openings, and every individual agent is protected by its team.
- Likewise is the map movement scary most of the time. Being able to recall just enough agents back to defend while simultaneously taking out strategic objectives of the human team. Also Blitz(caster) has noted earlier OpenAI's ability to focus on the winnable lanes and sacrifice the others, prioritizing well. (and exploiting map mechanics unknown to most pros until a few years ago)
- When the AI takes down a tier 1 tower, they seem to be very quick to take down the remaining tier 1 towers, instantly capitalizing on their map control advantage, and expanding it.
Some interesting things / bugs:
- Sniper bot throwing multiple spells on the same location right away (even though the damage doesn't stack) effectively simply wasting his mana and cooldowns, for no gain.
- Sniper using his ultimate ability to pressure the lower hp characters of the human team continously. Usually it's more often used more as a finisher, and might be an artifact that appears from the AI having access to an unkillable courier that ferries a lot of healing/mana regenerating items.
Some additional info learned from the interview of some of the devs:
- Incentivizing killing Roshan (a boss character in the middle of the map, which yields a one-time ressurection item for one player, after being killed) is done by varying roshan's HP down to a really low amount, making sure the AI experiences the upside to this. Otherwise it would require all 5 agents gathering there, expending their magical abilities and investing a lot before actually seeing a reward. (which is unlikely to happen)
- Game length in self-play sessions are above 60 minutes around 1% of the time.
Fogged (the human player) commented on this and said he messed up. If he had shift-queued the spell, or just used it immediately after blink, it would've landed [0].
Cancelling an initiation with instant spells (Lion hex, Rubick lift, etc) does happen frequently in high level human play as well, where you continuously pre-cast the spell, cancel, walk back slightly, repeat, on the out-of-range initiating enemy, to have the spell interrupt the initiation as soon as the initiating enemy blinks in to range. I do agree that the bots have a solid mechanical advantage, just pointing out that this specific scenario does frequently happen in human play as well (albeit not on every single initiation).
[0] https://www.reddit.com/r/DotA2/comments/94vdpm/openai_hex_wa...
This is the key - if we had machine learning techniques that allowed it to reason at the level of "what happens if I kill this?", we could explore more of the interesting state space more quickly. Perhaps there are advances in intrinsic reward systems that allow this.
Basically can I run the same sim on my laptop and watch them play? I can see some code on Github but dont know if the actually neural net data is available too.
If anyone knows the answer that will be great thanks
Once that is finished, you have a trained model which you can use by providing input and getting an output (with the weights frozen).
The training phase is very expensive computationally because you have to calculate the gradient of your loss function on potentially huge tensors.
The execution phase is not that expensive and commercial laptops will be likely able to run the model without any problems.
Just wanted to double check - can someone else verify this before I spend weeks seeing if I can get it to work?
A 1024 unit LSTM only takes up a few megabytes of memory, and the multiplications at runtime are O(N^2) and not O(N^2 M), because you don't have a minibatch of updates to run.
Even when reducing tagging accuracy to below human level, they still performed better.