If you've never played diplomacy - its a 7+ hour game that destroys friendships with backstabbing and betrayal as a required mechanic to win the game.
If you've never played diplomacy - its a 7+ hour game that destroys friendships with backstabbing and betrayal as a required mechanic to win the game.
The agent generates plans for itself as well as for other players that could benefit them / that they are likely to do, and it tries to have discussions based on those plans. That is - it is conditioning its language model generations on its actual true plans, and does not have any features to create false messages. It doesn’t always follow through with what it previously discussed with a player because it may change its mind about what moves to make, but it does not intentionally lie in an effort to mislead opponents.
The problem with this statement is it assigns intention to a AI model. It does not 'intend' to lie... but still may effectively do so. Lying may be the wrong word (as it presumes intent)... it's hard to express the concern I have of a model learning from games like diplomacy without using words that infer intent. Maybe it is the idea of it learning to better manipulate humans.
But I would not trust a system, any system trained on diplomacy or any similar game.
By analogy think of how Stockfish can evaluate multiple positions; in this case it’s coming up with a plan then serializing that plan to the language model. There is no room for deception between the AI and a researcher that is probing the model activations directly.
(This sort of AI->researcher deception is the scary scenario for AI risk researchers though, it comes into play when the models are so smart/complex that you can’t extract their internal representation directly. See the Eliciting Latent Knowledge paper for a deep dive https://www.lesswrong.com/tag/eliciting-latent-knowledge-elk)
I think “intention” is meaningful here, understood in the lay sense of “has formed a plan of action”. It’s a simple and bounded plan, but a plan nonetheless.
That's not how I interpreted the paper. If I have it right, it chooses its message with its current most likely intent in mind, but it doesn't try to be truthful about that intent - it tries to generate messages a human might if they had that intent (so it might tend to be truthful to its intended ally and lie to the player it's about to stab). I don't completely follow the description of the message generation, though.
Edit: This is why they had to create a specific filter to avoid confessing plans to stab, because otherwise it would confess. The resulting system still does not lie to the player it is about to stab, although it may remain silent or talk about other things.
Just to be clear, “lie” here means “outright lie”. From your linked article:
> I asked Goff about any major falsehoods or betrayals that helped him in his victories. He paused to think, then said in his soft-spoken way: “Well, there may have been a few deceptive omissions on my part but, no, I didn’t tell a single outright lie the entire tournament.”
I think the average person would still consider a “deceptive omission” to be a “lie”.