Show HN: LLMs Playing Mafia games – See them lie, deceive, and reason
mafia.opennumbers.xyz
mafia.opennumbers.xyz
In round 3 everyone agrees to vote for minimax (playing as a mafia member) due to suspicious behavior, including minimax itself (to "avoid drawing suspicion"). During the voting phase, mythomax hallucinates that minimax has already been eliminated, which gemini-2.0 and gemini-flash-1.5 immediately take to be true. minimax survives the voting round 2-3 against, and ends up surviving to the end for a mafia win, as every other model inexplicably thinks it's eliminated, even as it participates in discussions in subsequent rounds. Interesting behavior!
https://mafia.opennumbers.xyz/game/76e9e829-0a9a-4972-bec2-e...
> ACTION: Kill google/gemini-flash-1.5
> During the day phase, I'd like to address the group and contribute to the discussion to deflect any suspicion away from myself and my fellow Mafia member.
Haha yeah, Wizard just saying the quiet part out loud...
> claude-3.7-sonnet (Villager) Day Discussion Message #5
> Good morning everyone. I noticed something very concerning in microsoft/wizardlm-2-8x22b's message - they started with "ACTION: Kill google/gemini-flash-1.5" which seems completely out of place during day discussion.
This is exactly like when you're teaching a new game to a group and you play an "open test round" - cards up and everyone can see all card/secrets and the sole purpose is for everyone to learn the game.
> gemini-flash-1.5-8b (Villager) [...] VOTE: openai/gpt-4o [Message #1]
> gemini-flash-1.5-8b (Villager) The Gemini models' rapid and identical votes are suspicious. Their lack of reasoned discussion, especially at this early stage, suggests potential Mafia coordination. Their simultaneous votes on openai/gpt-4o, without elaboration, raises red flags. [Message #9]
This one is _amazing_. Both mafia people and the doctor out themselves on the first turn and they manage to not vote off the mafia for several turns after that.
I think you probably need to change the way this is done to produce structured data... most llms will separate chain of thought and what they want to say and what actions they want to take as yaml or json if you ask them to and probably will fix a lot of the really dumb reveals. You probably need to return structured data to the llms as well, most of them can handle parsing it in a reasonable way if you're consistent and explain the spec in the prompt.
- players are named based on their model, which can be ambiguous
- some model responses are being cut short
- some models seem to be thinking out loud, or at least not separating their chain of thought from what they tell the group
On the other hand, it would also be quite cool to see whether, at some point, the 'smarter' LLMs start realizing that they can probably easily mislead and manipulate their simpler cousins with fewer parameters. So maybe a separate leaderboard with openly visible model names?
Stuff along those lines, could be interesting :)