Iterated Prisoners Dilemma Strategies Dominate Any Evolutionary Opponent (2012)
researchgate.net
researchgate.net
You can write a perfect tic-tac-toe program with relatively simple rules, but an evolutionary strategy will wipe the floor at Go. This kind of modelling has value but real life is incredibly complicated, complicated strategies destroy simple ones as the rules of the game become more complex. Our brains are extremely expensive organs, and they're built that way for a reason. I think people are way trigger-happy extrapolating models like this to the real world.
The clearest example is how much energy we spend modelling other humans' thought processes. We spend so much because a simple permutation of tit-for-tat isn't sufficient - it will lose. It's an arms race to use complex strategies, an outcome which is specifically precluded in the model. I'd argue the fact observable reality is so divergent from the model means we should be skeptical of the applicability of the model.
Forgiving: always give, even to those who don't cooperate
Tit-for-tat: cooperates by default but does not reward defectors/freeloaders. This often the best strategy.
Interesting how this applies to software licenses.
Permissive licenses are clearly forgiving actors.
Copyleft/protective licenses are a gentler version of tit-for-tat.
Tit for tat, but also a small (random) chance of forgiving (giving a break) anyway. I believe the reason that won is that this allowed for recovery in the face of understandings and as long as the chance wasn't that large the cost also wasn't.
Interesting, if true, would the fist linkage between the Prisoner's Dilemma and the Utlimatum Game.
Normally in game theory, such statements are not seen as "credible", i.e. you assume the other person is bluffing and you go on to defect 100% of the time.
Ultimatum Game is about bargaining. "At what point are you willing to pay to punish someone for an unfair deal?"
Still processing the paper, but the implication here would be that you can link the two models, but the payoff to doing so critically depends on how much the dominant player is capable of modeling the other player's mind.
Aka smartest model/strategy/robot/person wins.
And in a pool of cooperative agents doing some "last-turn-defect" strategy while theoretically better than cooperate-always, is complicated with small payoff.
*Wherein the player does whatever their opponent did in the previous round.
[0] https://en.wikipedia.org/wiki/The_Evolution_of_Cooperation
The strategy is to submit lots of entrants to the tournament rather than just one, then play tit for tat against all opponents except those that are in the clique. The clique players then lose on purpose to a specifically-chosen player, maximizing that players score.
If you don't want to do it this way and instead care about maximizing the total scores of the clique as a whole, you can default to every clique member always cooperating.
Even in cases where the opponent is "blind", you can use a special pattern of bets to signal that the other player is a member of the clique and then play your win-maximizing strategy once the signal is detected.
This kind of thing has real-world implications. For example: imagine a poker game with ten players, all of the same skill level and without a rake. Nine of the players can collude (say, show each other their hidden cards and make decisions based on the shared information), making it significantly less likely than a 10% chance that the targeted player will win. I suspect this is already done to an extent in extremely high-stakes games (say, $1MM+).
The poker one already happens in real life all the time - if friends go to play at the same table, they often play differently than with a stranger.
See eg. https://www.nature.com/news/physicists-suggest-selfishness-c... which describes that political struggle going on for the last 40 years. This was eg countered by https://www.researchgate.net/publication/236189156_The_Evolu...
In most real world situations, the payoffs are much different than the PD ones.
Collaborate, and you may lose or win a little. Defect and for most cases, your payoff is 0.
Someone in a position of power may choose to set up a prisoner's dilemma - let's say, between you and your colleagues - to disadvantage you both, while giving an illusion of choice.