I wonder if this is an artifact of the training methodology: maybe if your team is very weak then your choices are also weaker, and reinforcement learning doesn't work as well?
I wonder if this is an artifact of the training methodology: maybe if your team is very weak then your choices are also weaker, and reinforcement learning doesn't work as well?
When the win percentage for Go AIs gets to around 5%, every action it can take results in a losing game so it can't make the difference between normal play and super strange moves anymore.
When every choice is really bad, humans tend to still go with their normal strategy and wait for their chance to turn things around, but bots assume the opponent is playing perfectly, so they act like their winrate is going to stay near zero no matter what they do.
Several levels of weak opponents should be used, with varying probabilities, to tune the AI’s robustness against real-world, imperfect competitors.
Taking a tier 2 tower nets everyone on the team 120 gold (a further 150+ gold goes to the hero who gets the last hit), and losing Sven probably gave the opposing team less than was gained.
Perhaps the AI simply placed more value on increasing the total net worth of the team than it valued saving the life of one of its core heroes. Additionally, there was no guarantee that he would have been able to escape, as Sven was deep on the enemy's side of the map, and there's a very real possibility that he could have been ganked from someone in the jungle had he attempted to retreat.
- Potential chance the enemy team would deny the tower before another friendly hero could take it (netting Sven's team 0 gold for the time spent whacking away at it)
- Map vision (removing a T2 often cuts a significant section of map awareness away, since the tower is no longer providing vision or protection)
- XP gains (Sven won't gain any XP while dead, nor from killing the tower)
- Creep equilibrium (this is less important, or at least thought about less often, later on in the game and past T1 towers, but might've been a factor in drawing the creep clash point to a particular location)
- Dictating team net worth averages (to some extent, if they predicted a loss in opponents forcing a teamfight or predicted a likely pickoff, gold lost could be minimized now by taking a death early, lowering the average net worth on the team).
Obviously, there are others and these can also be mixed and matched in various ways (e.g. cutting off map vision so they can more safely farm additional jungle creeps).
Not saying any of these aspects _were_ a part of the decision to trade Sven for a tower, but.. just wanted to include a few more subsurface aspects that _could_ be used in such a decision.
I'm not sure if the AI can surrender (I only managed to watch the first two games as it was rather late at night) but it might be a path to explore; having the AI give up if the game cannot be won anymore.
I wonder if this situation can be fixed by adding more randomness. For example, force AI'1 to be in a losing position to AI'2, but then suddenly switch the power level of AI'2 to be much weaker (where mistakes happen) so that AI'1 learns how to fight its way out of tough situations.
These are the kind of actions you specifically don't want to code in because you're throwing in human knowledge. You want the AI to learn by itself that using anti-invis when everyone is visible is a low-value move.
The purist in me was even mad that they had a hand-crafted evaluation function. (e.g. prefer gold, prefer taking towers, each given some arbitrary value)