The pirate game
en.wikipedia.org
en.wikipedia.org
Another problem which puzzles me even more is dividing a cake among three people by majority vote. Let's say Alice and Bob make an agreement where each of them gets half of the cake and Carol is left out. But that's unstable, because Carol can offer Bob a different agreement where Bob gets 60%, Carol gets 40%, and Alice is left out. Since Bob gets 10% more than before, he has an incentive to switch. But that's unstable as well, because now Alice can make a similar offer to Carol, etc. In the end there's no possible setup that everyone will stick with. As far as I know, this problem is still mostly open (though a few advances have been made).
To sum up, the idea of "perfectly rational" decision-making is surprisingly difficult to nail down. It's been definitively solved only for the case of two-player zero-sum games. When the game is not zero-sum, you get complications like bargaining and equilibrium selection, and when you have more than two players, you get an explosion of complexity without any clear-cut answers.
EDIT: As explained in the replies, this is not actually a zero-sum game.
I remember seeing this clip in 2012, and back then I proposed a solution based on correlated equilibria which still seems pretty good to me:
"We have two minutes to talk, right? I'm going to ask you to flip a coin (visibly to both of us) at the last possible moment, the exact second where we must cease talking. If the coin comes up heads, I promise I'll cooperate, you can just go ahead and claim the whole prize. If the coin comes up tails, I promise I'll defect. Please cooperate in this case, because you have nothing to gain by defecting, and anyway the arrangement is fair, isn't it?"
The verbal agreement "If you do this and I do that then I will split the money 50/50 with you after the show" is a verbal contract which (due to the cameras recording the event) is nevertheless indisputable. You are giving them a task to accomplish in exchange for a promised financial reward. They have all rights to sue you in court if they did what you asked and you didn't pay them for it.
What's really interesting is that he promises to split 50/50. In theory a rational actor who could commit to these things absolutely would say, "Look, I am pre-committing to steal, and I will give you $100 after the game if you choose split." Then it is still rational for the other guy to choose "split" (getting $100) rather than "steal" (getting 0). However most people, in study after study, will be insulted by the disparity of only getting 1% of the winning, and the insult is magnified by the camera: it is public shame. So in some sense being insulted by unfairness introduces out-of-band costs that also tend to pull you to splitting the money, and splitting it more fairly.
An anonymous, blind experiment would be more interesting, but less entertaining.
True story.
I don't follow. If pirate A assigns themself all the coins, the other pirates may as well vote against the plan - doing so risks nothing, and might result in them getting more than 0 coins if others happen to also vote against.
Maybe it's more helpful to think about such problems in terms of "how would I write a program to solve all problems in this class?", rather than "how would I act using common sense in this particular problem?" That often makes things clearer.
I would strongly challenge that assumption as having any realism (even for perfectly rational agents and no communication): if an agent decided to vote, he could assume all other perfectly rational agents (another usual assumption) would also take the same decision (by symmetry), so he can safely vote yes. This symmetry could only be broken if the agents have access to randomness.
If you take the time to think it though, you might well reinvent superrationality, updateless decision theory and other fascinating things. In fact, you'll quickly get to questions that I have no idea how to solve!
In the case of Prisoner's Dilemma, if the cooperate/defect outcome had sufficiently large cooperation bonus for fixed other payoffs (large enough T>>R in wikipedia's notation), then with my assumptions it's clear that each prisoner should flip a coin to decide, and each one gets ~T/4 expected payoff.
But the most glaring problem is of course assuming every player is perfectly rational. Personally I'd only assume that if your players are all game theorists with enough of time and paper :) I probably wouldn't even assume myself as rational.
In conclusion, I believe equating maximin with optimal play/perfect rationality is misguided, but maximin is a good safe bet.
Like someone else said, you seem to be speaking in jargon. "Superrationality", "updateless decision theory", none of this matters in order to understand the Pirate Game. The rules of the game are very simple, maybe the pirates thing is throwing you off? That's just a theme for the puzzle (people aren't rational like this, and of course pirates aren't, and the whole situation is extremely artificial and unrealistic). The non-intuitive solution is derived from the rules as stated without any need for additional theory.
The point made was that your 'equilibrium' is not valid because a pirate voting that way decreases their expected return. I don't see how you need a different 'equilibrium concept' for your proposed solution, when the basic assumption of rationality excludes it.
Not trying to be mean, but your response seems obfuscatory.
Ideas like Nash equilibria and subgame-perfect equilibria are the only known formalizations of rational behavior in multiplayer games. I'm not obfuscating, I just don't know any other kind of math that would work...
If everyone knows what everyone else is going to do, then neither option is better or worse, sure. But where did you get this perfect knowledge from? Certainly you couldn't deduce your situation.
And why is 'no' your default choice: this isn't a repeated game. You're trying to decide what to vote, you don't have a mind to change, at this stage. And neither does anyone else. Therefore, if everyone knows I can choose to vote either way, your situation is irrational.
You're objecting to a problem that isn't the one being considered.
It may well be that there is a compelling mathematical argument you're trying to evoke, but at the moment you're not actually putting it forward. You're name-dropping, but not explaining. Well, not in a way I can follow, anyway.
> In game theory, the Nash equilibrium is a solution concept of a non-cooperative game involving two or more players, in which each player is assumed to know the equilibrium strategies of the other players, and no player has anything to gain by changing only their own strategy.
Many people have tried to define mathematically what "rational behavior" means when there are multiple players. Nash equilibrium is the one idea that stayed. It's the first chapter of every game theory textbook. Since the 1950s, literally all publications about non-cooperative games have been based on Nash equilibrium. To understand why, you'd have to look at the history of the field, or you could take my word for it. If it's any consolation, it was very unintuitive for me at first, too.
"If each player has chosen a strategy and no player can benefit by changing strategies while the other players keep theirs unchanged..."
which is not the situation in this case.
Strategies for infinite games are not necessarily optimal if the game is finite, and vice versa.
That's a pretty unusual setup, because game theory usually just gives you a utility function over outcomes and lets you derive strategic behavior from that, without fiddling with individual actions. Though I guess most people solving the problem have tried to merge the two utility functions somehow? That's probably the root of the disagreement in this thread...
Equilibrium meaning 'if everyone votes this way, there is no benefit for any individual to vote differently' - that doesn't apply when the pirates are trying to determine what to vote for the one and only time they'll ever vote on this question, not knowing what anyone else will be voting.
Which is why you've been unable to express your disagreement with the puzzle without resorting to language of 'changing' votes, or assuming that you know what other people are going to vote. None of which are reasonable assumptions here.
Right, so it meets the definition of an equilibrium.
> but is this really interesting for this game, i.e. is it a viable outcome of it despite being stable once reached?
In "reality" no. But the Nash Equilibrium formalism gives us no way of determining which of the two possible equilibria (this one and the usual solution) will happen in practice. So you need some more powerful idea to solve that (see references to "superrationality" and "trembling hand" elsewhere in this thread), and none of those theories is mathematically complete.
A pirate will vote YES if offered 0 when the alternative is getting killed in one of the following turns. Pirates always prefer to live first, and only then they prefer to maximize their money.
edit: to clarify, the only pirate that would have to consider being killed is the one making the offer. thus yes, offering yourself 0 is possible and in fact seems to be a requirement for the (one of the) stable solutions. But if I'm not most senior, I would never accept 0.
Theoretically, a pirate could accept a seemingly inconvenient proposal if he evaluates that rejecting it would result in himself getting killed further down the line, when it's his own turn to propose a split.
You're right that this never happens in the 5-pirate game as stated (I'm not sure about the general N-pirate, C-coins game). The pirates who are offered 0 coins reject the split, but they don't get to alter the result anyway.
Equilibria just seem to me to be the wrong tool for the job, you have to (as you have) invent a related infinite game with additional axioms and expectations before that kind of analysis kicks in.
Check your responses. If you can't describe what you're trying to say without mentioning anyone 'changing' their strategy, or assuming that a strategy is default, then you've got the wrong framework, I'd suggest (modulo my admitted lack of specific expertise).
> Third, each pirate would prefer to throw another overboard, if all other results would otherwise be equal.
If a pirate found themself in a situation where they thought their vote wouldn't change the outcome -- i.e. all other results would otherwise be equal -- they would vote to murder the proposer. In the situation you describe, every pirate would find themselves in that situation, and would thus vote for blood.
Wouldn't Alice, Bob and Carol realize that then and agree to divide the cake evenly since it's the only solution that is fair to everybody?
For a single cake, offering a fair split is just as hard to predict success as any other split. It also assumes there are no other signals, which is why it's not very insightful to begin with.
Is there a name for this problem? I can only find the fair cake cutting problem, which has been solved for 3 people.
Or more specifically, https://en.wikipedia.org/wiki/Envy-free_cake-cutting
On a tangential note, something that's always bothered me about write-ups of "envy-free" division is that they assume a couple of things I don't think are true:
- That someone cutting a cake always executes exactly the cut they planned.
- That once the cake has been divided, no one will experience envy because their piece is worth, to them, at least their "fair share" of the total.
Take the classic example of envy-free division of one cake among two people. The solution is simple enough that most people understand it instantly: person A divides the cake into two portions however they please, and person B then selects one of those portions, A receiving the other. Everyone will agree that this is intuitively fair. But it's certainly not fair for either of the two reasons I dispute -- A might still envy the portion of the cake that went to B, imagining a different world where the entire cake was theirs... or A might slip up with the knife and create one portion which is obviously better than the other portion, which B then chooses.
Rather, I think the benefits of envy-free division come from two other points:
- The parties involved agree beforehand that the process is fair. This requires them to be able to understand it, which is easy for the two-person division and quite difficult for more than two.
- Instead of assuming that division goes off without a hitch, we can observe that if you receive a portion you're unhappy with, it's because you messed up. The blame falls on you rather than someone else.
A group of superrational players may choose an outcome which is not even a NE! For example, two superrational players playing the prisoner's dilemma game will choose to cooperate.
Conversely, Nash equilibria are no longer "stable" in the original sense that there is no incentive for any player to unilaterally deviate. The concept of "unilaterally deviating" doesn't even make sense, because a group of superrational players will all choose the same strategy. The superrationality allows them to cooperatively deviate.
To demonstrate, I consider the smallest case for which your equilibrium strategy differs from the wiki pages' strategy, that is n=3. Suppose there are 3 pirates, A, B and C. Suppose the first pirate proposes that he gets all the coins. Then the state where everyone votes yes is a Nash equilibrium because no one has an incentive to unilaterally. However I claim that a group of supperrational players will vote to kill A.
Here is the proof of my claim (caveat: I don't think this is actually a valid analysis - see next paragraph. however it's the same level of rigor as the wiki page and demonstrates how B and C can cooperate to kill A, "escaping" the NE of both voting Y). C will vote No because accepting the proposal is the worst possible outcome for him (he gets 0 gold and no one dies; note that it is not possible for C to die), so strategies where he votes No in this round dominate strategies where he votes Yes in this round. B knows that C is thinking this. Hence he will vote No.
Explanation of caveat: actually, the nonstrict dominance proof in the above paragraph suggests that C does not know how B will vote and the dominance is nonstrict because of this unknown. However, since B and C are superrational, there is no such thing as not knowing how someone would vote; C knows that B will vote No, so the dominance is strict. However this argument seems circular (even though, because of superrationality, I think it really isn't).
Unrelated point: If we ammend the rules of the pirate game to include another tiebreaker with lowest priority: that all pirates prefer to vote "no", then your situation is no longer even a Nash equilibrium (much less a subgame-perfect NE). Then how many NEs are there in the ammended game?
——————————————————————
Imagine a 100 rounds of ultimatum games. In each round, Alice proposes a split of $1 and Bob can either accept or reject. If Bob rejects, the $1 is burned in a fire. How much money will each end up with?
Lets work backwards, as we do in the pirate game.
In ROUND100, Alice knows Bob will take whatever non-zero offer she makes. She can offer him 1 cent and Bob will agree.
In ROUND99, Alice knows that Bob can will accept whatever non-zero offer she makes in ROUND100, so she can make whatever non-zero offer she wants this round without fear of repercussion. She offers him 1 cent and knows he will agree.
Continuing to work backwards, we find that in all 100 rounds, Alice takes 99 cents and offers Bob 1 cent. Alice ends up with $99 and Bob ends up with $1.
——————————————————————
This is obviously (and experimentally proven) not what would happen in real life. In experimental outcomes, Bob rejects offers that are too low to broadcast that he is “irrational”. As a result, Bob is able to negotiate a much higher outcome (around 40%).
So who is more rational? The "rational" Bob who gets $1, or the “irrational” Bob who gets $40?
I might be mistaken but I think in terms of economics that this is literally the definition of rationality.
Your experimental example is with real people. Real people are not rational actors (as a rational actor is defined in economics). This is considered by many to be a major flaw in traditional economics.
On the flip side, this is considered by many traditional economists to be a major flaw in real people.
The insight obtained from from comparing the computer-like rationality to what would happen in the real world is truly fascinating to me.
Also, your example of Bob negotiating fits in very well with the theoretical computer-like rationality. Alice's offer of 1c is equivalent to an offer of 0c nominal in the hypothetical case.
The real-life 40% is more a commentary of utility of money as opposed to decision making. I think that you're conflating concepts.
The real-life 40% is more a commentary of utility of money as opposed to decision making.
Only real-life humans make decisions though. If there is no implied utility of money then Bob might as well reject all offers. And it does seem that the only important question is "what would real humans do?" We can set up a simulation that strictly adheres to a simple set of rules and watch how it plays out, but what is that telling us?
It's a logic puzzle. We learned that the intuitive solution of having the first proposal be "I don't want any money, I just want to live" is too conservative, and that there is a way better solution for the first pirate.
We didn't gain any deep insight in how to split money between real people, since those cases seldom involve pure, unemotional logic.
In a game where "Bob" is continually offered one cent, if he burns the other player who is trying to get the other 99 cents, he's signaling that they'll have to cut him a better deal. The other player would be irrational to continue offering the 99/1 split after it had been previously rejected.
For the other player, even (especially!) if it's a computer, the rational choice is to keep upping Bob's percentage until it finds a number he'll accept. So you're right that 40% only speaks to a specific utility, but in any case, the only rational number is whatever number Bob (and the first player) will accept.
Wait, how does she know that? Can you explain why you consider that the most rational choice for Bob to make? As experiment shows, this is not the optimal strategy, so I don't see why you would call it the "rational" one.
Also, keep in mind that you're describing an iterative game (where the same people play the same game with the same circumstances repeatedly) and the pirate game is non-iterative, since the players and circumstances change between rounds. As you note, iterative games are much more complicated and the sort of induction used in the pirate game doesn't really work.
It's rational because if Bob accepts an offer for a positive value of money, even if it's a cent, he gains value from it. Rejecting it just throws that cent away, which he could have had in his pocket, it's irrational to throw away this value.
There's a difference between rational and optimal for a reason though, but you need to feel like you're getting ripped off for that to come into play, you need to feel like you'd rather someone else lose $0.99 than you gain $0.01.
I agree that if it was a one off game, then that would make sense, but it's an iterated game.
Archive.org have a mirror:
http://web.archive.org/web/20081218012123/http://history.byu...
The usual format is that they play a game each episode, and the winner and another player of their choice are guaranteed to go to the next round. The loser of the main game then plays a one-on-one game with another player of their choice, and the loser of that game is eliminated. This creates an interesting metagame surrounding the main games.
One thing that makes it work particularly well is that the participants all seem to be relatively pleasant people. There were alliances and sudden-but-inevitable betrayal, but no one was ever unpleasant about it. I'm not sure if it's a cultural thing, or that they're all notionally "geniuses" (talented in some field) rather than random people or celebrities, but that atmosphere really made a good premise into an excellent show.
Each such problem I've read about assumes that we can assume causality while reasoning about hypothetical, but strangely, if we let go of that assumptions, then we can arrive at different answers which hinge on otherwise surprising behaviour. I relate this idea to the fact that in logical reasoning, all true statements are, once proved, held to be simultaneously true. That is, given if A then B, with A being true, we don't hold B to be true 'after' A, but to have been always true, given A.
In the pirate problem, limiting ourselves to three pirates to shorten my explanation, we end up with the split being: 99 to C, 0 to D, 1 to E. This is because we assume that if only D and E were left, D would keep 100 to himself. Now given that distribution (99,0,1), D should now change his hypothetical proposal to (0,98,2). If we assume that C is not a 'non-causality' believer, but both D and E to be, then E would vote no to (99,0,1) and D would do the (0,98,2) split. You may argue that D could then do a (0,100,0) split, but that's not how a non-causality believer MUST act, because he knows that to get to that point, he must know thet E can logically know that he will do this split. This can be justified by arguing that when a pirates survives, he will enter such gold-splitting game later on. But my argument is subtler than this and doesn't require it. It basically become this: all such pirates, posited to be perfectly logical, are interchangeable. Thus they must all reason in the same fashion. Thus, my argument is that true pure logical minds see that the true way to maximize their gold profit is to hold a world view that maximize their profit, even if causality is discarded. Thus both D and E know that tehy can maximize their profit by discarding their belief in causality. That is how D ends up proposing (0,98,2) and E accepts it.
Of course, I ended up there by assuming C was not a believer, but given my argument, C must also be ready to throw causality out of the window, otherwise he will end up dead. I believe my argument ends up splitting (49,51,0), but I'm not sure. Once causality is throw out, it's hard to tell, but intuitively, with three pirates, those voting must be given almost equal gold and the remaining pirates must not have a majority.
Anyone interested in a more complex and realistic examination of the economics of pirating should consider Peter Leeson's book, The Invisible Hook: The Hidden Economics of Pirates.
I would also recommend Leeson's Anarchy Unbound: Why Self-Governance Works Better Than You Think for some interesting applications of game theory in more realistic historical contexts.
Much like a two-body system is useful for explaining gravity, even though there's virtually no place in the real world where that particular example would be useful.
When models utilize more realistic assumptions, they can be incredibly powerful tools for understanding past behavior and predicting future behavior. Historical evaluation is particularly powerful when game theory is utilized with the analytical narrative form of analysis.
Quite an entertaining problem, skip the result and try it out yourself, the sense of accomplishment is fulfilling, especially if your mood is shaky.
If I were C I would have voted no on A's proposal and try to make a deal with D and E. The idea is that knowing if A were thrown overboard then B might be more apt to provide a more beneficial coin sharing program in an effort to save his life. After all, he just watched the previous guy get tossed overboard and die. The most I could lose would be 1 coin if B decided to try A's method on D and E and they went along with it.
The key is that I would have already negotiated a deal with D and E, if A or B doesn't share the wealth fairly then we vote to toss them. Once I'm in charge, I'll share the coins equally as possible.
I read that as "an actual proposal, by the leader" not as a "theoretical future proposal, iff that negotiator becomes the leader" and I believe that's the intended reading.
Thus, D and E cannot trust that you'll hold your word to give them 33 coins each provided you get to be in charge.
But yes, I glossed over that part apparently.
"I have the turn and I propose we split this way": trusted.
"I don't have the turn yet, but when it's my turn, I promise I'll propose this / kill him / won't kill you": not trusted.
Anything else cannot be trusted. Because the pirates are completely rational (think robots), they not only don't trust each other, they also know the rest don't trust them either. So no pirate will attempt any kind of deal.
In a one-off game a rational player C will negotiate and then propose a 99-0-1 distribution.
When A is in charge this affects....
When B is in charge this affects...
Maybe that's recursively?
solve(pirates) {..... solve(mutatedPirates) }
Note to self: don't get thrown overboard.
If pirates, all things being equal, prefer violence to peace (which is implied) then E getting 1 from A is less preferable than killing A and B and getting 1 from C.
Or is it that E knows that B will be offering next, and B will offer him 0, so it makes sense for him to accept A's offer of 1?
Yes. That's why it is sufficient to offer E: 1.
According to the pirate's code, a bribe is a binding contract, but only to the extent that breaching the terms of the contract results in returning the amount of the bribe after the breach occurs. Otherwise, the bribed pirate is thrown overboard. A pirate may therefore spend the bribe money before reneging, if reneging will still yield enough money to pay back the bribe afterward.
Pirate P[x] has savings s[x], and s[x] > s[x+1].
The degenerate case where n=1 is easy to deduce. P[1] proposes {pool}, votes yes, and wins.
The case where n=2 is also easy. P[1] proposes {pool, 0}, votes yes, and wins.
When n=3, it gets more complicated. P[1] needs one more vote to win. If P[2] is able to propose a split, he will be proposing {0, pool + s[1], 0}. So the cost of P[2]'s vote is at least pool + s[1] + 1, which is normally impossible to achieve for P[1]. But since P[2] does not know how much s[1] is, other than that it is more than s[2], P[1] might be able to risk it. But since P[1] doesn't know s[2], bluffing a lower amount for s[1] is risky, as if the bluff amount is lower than s[2] + 1, then P[2] will immediately know to vote no. The cost of P[3]'s vote is at least 1. P[1] and P[2] will therefore be bidding competitively for P[3]'s vote. P[1] has the choice of offering value to P[3] publicly in the proposal, or secretly as a bribe. Either way, P[1] is likely to propose 0 for P[2].
So if P[1] proposes { pool, 0, 0 }, and bribes P[3] to vote yes with s[1], P[2] might try bribing P[3] with any amount from 1 to s[2] to vote no, with no effect. The net effect is { +pool -s[1], 0, +s[1] }.
Remember also that any bribe P[1] pays to P[2] could be added to a bribe from P[2] to P[3]. If P[1] bribes 1 to P[2] to vote yes, and s[1]-1 to P[3] to vote yes, and P[2] bribes s[2]+1 to P[3] to vote no, then if P[2] and P[3] both vote no, P[1] is thrown overboard for losing the vote, and P[2] is thrown overboard for reneging on a bribe and not having the coin to pay it back, so P[3] gets everything. If P[3] suspects that P[1] may have bribed P[2] to vote yes, and knows that P[2] bribed him to vote no--which would be useless unless P[2] himself intended to vote no--then P[3] may vote no on the possibility of getting pool + s[1] + s[2], which would otherwise be impossible for him. So knowing this, P[1] may be confident that any bribe to P[2] to vote yes would not reasonably be re-bribed to P[3] to vote no.
So P[1] could bribe 1 to P[2] to vote yes, tell P[3] that he had bribed P[2], tell P[2] that he had told P[3], and then propose {pool - 1, 0, 1}. It would be useless for P[2] to bribe P[3] unless he intended to reneg, he can't add P[1]'s bribe to his own, and he knows that s[1] - 1 >= s[2]. With the 1 from the pool, he could not match such a theoretical bribe anyway. So the net effect is now { +pool -2, 1, 1 }.
I'm not exactly sure if that is a stable solution or not.
I think maybe that P[3] might be able to get more by bribing P[1] to vote no, or P[2] to vote yes.
Oh, get a job? Just get a job? Why don't I strap on my job helmet and squeeze down into a job cannon and fire off into job land, where jobs grow on little jobbies?!
Why are you so disturbed about the ridiculousness of my comment, yet quite happy to take the ridiculousness of the situation's premise?
If ANY of the respondedants and downvoters to my comment actually wanted a touch of realism, they'd understand that I was pointing out the ridiculousness of 'perfectly rational actors' in these games.
Basically, they had the same advantages as pirates near Somalia and Indonesia have today - very few military ships, and merchants are unwilling to fight because it risks the entire crew being massacred.
Why do you think rationality and survival are specifically at odds with being a pirate?
I'm actually trying to inject some humanity into the problem, highlighting (as others have) that we're talking about humans here, not logic gates. If the pirates were so extreme in their rationality, then they wouldn't be pirates in the first place. Talk of such grandiose concepts like self-determination seems out of place when you have 80% of the crew being satisfied with a measly two coins.
"If the pirates were that strictly rational and interested
in their own survival, they wouldn't be pirates"
to which I replied "Why do you think rationality and survival are
specifically at odds with being a pirate?"
Your comment is distilled down to if strictly rational and & interested in survival, then not pirates
So my question asking why you think this is not following incorrect logic, let alone a strawman -- if you're going to be pedantic about rhetoric then at least know what fallacy you're accusing of me means. You can't just say 'thats a strawman!' when someone drills down on your argument.I mean, I even italicised the word 'strictly' in my reply, to clearly indicate the important bit of my comment that you blithely removed to create your strawman. You didn't 'drill down' my argument, you twisted it to say something I never said. It's like if I said "if a person inhales too much water, they can drown" and you replied "why do you think water is at odds with survival?".
Hell, even in this new "distillation" of yours, the word 'strictly' is used. If you're a fan of rhetoric, then get on board with the meanings of words. And if you want to be pedantic, I didn't just say "that's a strawman", I pointed out why it was.