CICERO: An AI agent that negotiates, persuades, and cooperates with people
ai.facebook.com
ai.facebook.com
Code: https://github.com/facebookresearch/diplomacy_cicero
Site: https://ai.facebook.com/research/cicero/
Expert player vs. Cicero AI: https://www.youtube.com/watch?v=u5192bvUS7k
RFP: https://ai.facebook.com/research/request-for-proposal/toward...
The most interesting anecdote I heard from the team: "during the tournament dozens of human players never even suspected they were playing against a bot even though we played dozens of games online."
"Screenshots of the stab below.
The human player said: "The bot is supposed to never lie [...] I doubt this was the case here" "I was definitely caught more off guard as a result of this message; I knew the bot doesn't lie, so I thought the stab wouldn't happen." "
"I'd like the researchers involved to say quite a bit more about "A.3 Manipulation"
What are possible prevention, detection & mitigation steps?
What are the possible use cases? What are the benefits/downsides of them? Has Meta considered developing products based on this?" -- Haydn Belfield, a Cambridge University researcher who focuses on the security implications of artificial intelligence (AI).
https://twitter.com/HaydnBelfield/status/1595168102924402688
On the other hand, the bot has no concept whatsoever of keeping its words. After saying words, it is free to change its mind about what moves to play, motivated from, for example, messages from other players.
> On the other hand, the bot has no concept whatsoever of keeping its words. After saying words, it is free to change its mind [snip]
Reminds of that one Asimov story about the robot who had a different interpretation of the first law of robotics. If my very hazy memory is right, the idea was that the robot could put a person in danger if it knew that it had the ability to prevent any damage from happening, but once it caused the danger, it could choose not to act and allow the person to come to harm.
I might be remembering this incorrectly, it's been a very long time since I read the story, but that was the first thought that came to mind when reading your comment :).
It goes with the zeitgeist to argue for what makes the life of big tech companies hard, but they are big enough that they can afford things like that. It's smaller companies and academics that would end up not being able to innovate as much
Go down that road and you end up with an IRB evaluation requires for an A/B test that changes the color of a button
Creating an AI to lie seems like the Wrong Path.
Zuckerberg should shut this down ASAP. If that's the case.
Randomly generate a city full of people. Make a few dozen of them the important NPCs. Give them situations and goals, problems they need to solve and potential ways to solve them. Certain NPC's goals are opposite others'. Then drop the player into that world and have the 'quests' the player is performing be generated based on the NPCs needing their help.
Updates wouldn't be adding new hand-written stories, it would be adding more complexity, more goals, more problems, more things that can be, and the story would generate itself.
Done right, this would be incredible.
I want to run a game of DnD using a heavily modded Crusader Kings or Civ to simulate the world they are playing in. Just set up the starting conditions, like the various kingdoms, their methods of succession, their relationships and blood feuds with other kingdoms, the internal treacherous machinations and possible rebellions brewing, wars between kingdoms, etc, and then just let it run and use that to generate the overarching plot events of the setting while somehow translating the players actions as some kind of input into the simulated world so they can affect it as well.
Most likely wrong if you compare to how most other development in engineering happened.
From my point of view, the most likely path towards some form of weakly general intelligence at this point is emergence. We keep working on these narrow problems and from the broadening networks at some point we end up with something indistinguishable from a general ai inadvertently.
Planes don’t fly like birds. There is very little reason for what you would call a “general intelligence” ai to develop in a way which mimics our own intelligence. It might but I would find that more notable that if it did not.
I'm fairly certain you cannot rigourously define "proper true AI", and I'm also fairly certain you don't mechanistically understand how humans think. This raises the question: where does your confidence that our current path is not pretty close to "proper true AI" and/or how humans think come from?
Machine learning models have written multiple scholarly papers that have been accepted to prestigious journals. Find me a rabbit that did that.
/s
Abstract: Despite much progress in training AI systems to imitate human language, building agents that use language to communicate intentionally with humans in interactive environments remains a major challenge. We introduce CICERO, the first AI agent to achieve human-level performance in Diplomacy, a strategy game involving both cooperation and competition that emphasizes natural language negotiation and tactical coordination between seven players. CICERO integrates a language model with planning and reinforcement learn- ing algorithms by inferring players’ beliefs and intentions from its conversations and generating dialogue in pursuit of its plans. Across 40 games of an anonymous online Diplomacy league, CICERO achieved more than double the average score of the human players and ranked in the top 10% of participants who played more than one game.
> Cicero ranked in the top 10% of participants who played more than one game and 2 nd out of 19 participants in the league that played 5 or more games.
> As part of the league, Cicero participated in an 8-game tournament involving 21 participants, 6 of whom played at least 5 games. Participants could play a maximum of 6 games with their rank determined by the average of their best 3 games. Cicero placed 1st in this tournament.
This bit seems a little more impressive I think. Being in the top 10% of people who’ve played at least two games might leave a lot of bad players to beat up on. Winning a tournament might (?) mean you have to beat at least a couple players who understand the thing.
It is sort of funny to think about — anyone who gets really legitimately good at anything competitive goes through multiple rounds of being the best in their social group, and then moving on from that group to a new one that is comprised of people who were the top-tier of that previous level. It isn’t obvious to me where on that informal ladder this tournament was.
But anyway, maybe the AI will follow the trajectory of chess AIs and quickly race away from human competition.
Similarly, in Go it seems unlikely that perfect play could overcome a nine-stone handicap (again, I could be wrong, I'm not remotely a dan-level player).
All to say, it seems likely that Diplomacy is a game where the difference between "the best human play" and "the best possible play" is much larger than either Go or Chess.
I figured 9 was too much, but I had no idea what number less than 9 to pick, so I stuck with what I thought was a near-certain upper limit.
I once played with 9 stones against a dan level player. It was close...until it wasn't :-)
Definitely, Diplomacy in general is substantially understudied compared to Go or Chess (largely because it's a tiny community). You can play for less than a year and get to top-level performance and much of the established wisdom/strategy of players is fairly bad.
Even the best Diplomacy players are only scratching the performance of how good someone could be.
Diplomacy is a bit of a choose your own adventure game too. Like there's an objective criteria (average SCs at the agreed end of game) but the human tendency is to try and win individual games. Humans will often choose to play sub-optimal strategies for better entertainment value.
I think the real accomplishment here is the ability to fool humans into thinking their not playing a bot. That's an impressive thing to do even these days.
... Isn't this the bad path of AI research? An unbeatable and utterly convincing conversational AI that knows exactly what you want to achieve, and then completely nullifies your attempts at reaching your goals while simultaneously achieving it's own?
So far all this has been with one player, amid others, no collusion.
You’re going to be surrounded very shortly by sleeper bots, including on HN. Relying on dang and others to root out bots will be futile. A swarm could easily collude to downvote people or get them ostrasized by their own friend group, as we have already seen when it came to crypto, metoo, BLM, lockdowns, vaccines and now Ukraine.
It’s really not hard for a bot swarm to completely exploit society in these and many more ways, and we are not ready for it. By the time the botswarms arrive online, it’ll be too late to do anything.
Update: LinkedIn already has a huge problem of fake profiles applying for jobs and offering jobs. And Twitter is overrun with bots. But soon, bot-written articles will be out-shared by your own friends rather than that “hack liberal/establishment/hasbeen paper” NYTimes.com
The animatrix had a good sort of storyline on this, but it involved a lot of unrealistic violence
You could for example limit users to those that log in with electronic id's issued by a government or other organisation that you trust to assert that the user is human and then force real names or a single user name for that e-id.
This already happened in other areas of life. Both fathers and mothers now neglect their own children and elderly parents so they can work for corporations. They often prefer this and find meaning in climbing the corporate ladder. Eventually, their own labor will be rendered obsolete, but for now they're in a race to the bottom to work harder and neglect their family even more. They even stick them in nursing homes.
Also, you no longer want to ask people for directions, you use Google Maps. You no longer ask your parents, teachers or libraries when you can just look it up online with no judgment.
Finally, look at industries like Wall Street trading. It used to be a bunch of guys in a pit. You'd call up your broker or whatever. Now everything is automated with bots. Everyone prefers bots. They make up the bulk of trading with real capital. These bots are are working for corporations, which employ less and less humans.
So the present is already a bunch of corporations owning bots and bots creating content for other bots. In the finance industry. Now how different is a bunch of text generation online? I think the human contributions will be vanishingly small in most communities.
The question is ... what is this all for? Dropping demand for human services is a byproduct of making things more and more efficient...
This tech is not a quality of life enhancer at all, it's our competition as a species...but then again if we are really this stupid maybe we (well, you lot) deserve to die at your robot overlords hands.
> ... by inferring players’ beliefs and intentions from its conversations and generating dialogue in pursuit of its plans.
As opposed to what? Not inferring intentions, or generating dialogue against its own plans, or at random?
It's just doing what a person does. It understands the other person, and says things to further its goals.
I can't help but wonder at what point I will inhabit a world indistinguishable from that of a paranoid schizophrenic. Will I even notice? And if I do, will anyone else? When we become as slow as trees to digital arborists, what will become of us? Will they domesticate us? Will they deforest us as we did Europe amd the Near East? Quo vadis, Domine?
My wife has schizophrenia, which is well under control with medication. But about every year or two, when she wakes me up in the night to tell me she's panicking because someone hacked her smartphone and laptop etc., I know we have to adjust her medication for a few weeks. It was scary at first, but now we know the drill: a night or two without much sleep, no big problem. Of course I check smartphone, laptop etc. You never can be sure, can't you? Especially not if you've already been in trouble with credit card fraud twice.
I once asked her psychiatrist what he would say to his patients who believe they are being monitored in these times of Snowden and Co. He said it didn't make his job easier, but he would calm them down and adjust the medication. After all, he knew that his patients were sick.
So what do you think they or their owners/masters will do when All is Watched Over by Machines of Loving Grace in our/their future/reality?
My prediction is 2029
Eagle Eye was a meh movie but the concept that people's own friends could be made to ostracize them and coerce them to do things, is a major concept in that movie.
You don't need violence to make it happen. You just need a swarm of AI bots to coordinate to reputationally outcompete others on all networks that matter. This is an optimization problem with a measurable metric (reputation). Bots can simulate the game among themselves and evolve strategies that far outclass all humans. It'll be like individual Go stones playing against AlphaGo placing stones.
You won't see it coming. The thing is, once they amass all that reputation by acting "relatively normal", you'll see so many kinds of stuff you won't believe. Your entire world could be turned upside down for very cheap. Reputational attacks were already published by NSA: https://www.techdirt.com/2014/02/25/new-snowden-doc-reveals-...
And this is just humans doing it. Bots can do this 24/7 at scale to pretty much everybody, and gradually over a few years destroy any sort of semblance of societal discourse if they wanted. They'd probably reshape it to suit the whims of whoever runs the botswarm, though.
Can you accurately pigeonhole that thought onto one of these? I'm curious how unique a storyline it is.
https://en.wikipedia.org/wiki/The_Thirty-Six_Dramatic_Situat...
Have you seen the TV tropes site? Chock full of awesome stuff :)
Silicon based AI isn't the only form of AI that can get up to mischief.
Incarnate technology deceiving humans is a different domain entirely. Human motivations are ultimately comprehensible by other humans.
But what do I know of the motivations of (for example) the Bible? What happens when the living incarnation of a holy text can literally speak for itself? Or when the average believer thinks it can?
Ultimately this may all amount to the same status quo, but seeing what cognitive distortions have come along with literacy, newspapers, radio, TV, and now the Internet, I have to continually ask, will I personally be able to maintain skepticism when the full brunt of an AI and its organs is suggesting faith otherwise?
You can call me an alarmist or melodramatic if you wish, but it should give everyone pause that the delusions of paranoid schizophrenics from the late 20th century are now basically indistinguishable from emerging popular technologies and their downstream effects.
And despite a substantial subset of the population knowing this, we continue to do nothing to address it - if anything, more people are devoted to giving people even more powers to deceive at massive scale.
> Incarnate technology deceiving humans is a different domain entirely. Human motivations are ultimately comprehensible by other humans.
Whether they are accurately comprehensible is another matter though.
> But what do I know of the motivations of (for example) the Bible? What happens when the living incarnation of a holy text can literally speak for itself? Or when the average believer thinks it can?
Likely: mostly nothing. Thus, the subconscious mind steps in and generates reality to fill the void.
> Ultimately this may all amount to the same status quo, but seeing what cognitive distortions have come along with literacy, newspapers, radio, TV, and now the Internet, I have to continually ask, will I personally be able to maintain skepticism when the full brunt of an AI and its organs is suggesting faith otherwise?
Do the laws of physics prevent you?
If not, then what? And, have you inquired into there is any pre-existing methodologies for dealing with this phenomenon?
> You can call me an alarmist or melodramatic if you wish, but it should give everyone pause that the delusions of paranoid schizophrenics from the late 20th century are now basically indistinguishable from emerging popular technologies and their downstream effects.
I am far more worried about the delusions of Normies, as they are 95%+ of society and are for the most part "driving the bus", whereas schizophrenics account for a small percentage, and tend to not be assigned many responsibilities.
One of us is more correct than the other - how might we go about accurately determining which of us that is?
What difference does it make that it's a computer doing it?
https://en.wikipedia.org/wiki/How_to_Win_Friends_and_Influen...
1. Because humans have an understood limit on intelligence.
2. Because we have systems in place to keep humans in check.
3. Because humans have a distinct physical location and thus stricter limits on their direct and indirect influence.
4. Because humans can't copy themselves.
Blog: https://ai.facebook.com/blog/cicero-ai-negotiates-persuades-...
Site: https://ai.facebook.com/research/cicero/
Expert player vs. Cicero AI: https://www.youtube.com/watch?v=u5192bvUS7k
RFP: https://ai.facebook.com/research/request-for-proposal/toward...
You don't win Diplomacy by staying honest, you win by choosing the right moment to make the move and burn that trust. It's the 5% dishonesty that matters.
I can't attribute that the parent - I don't know the individual - but I expected a chorus of such responses. I feel like I can see them squirming in their seats ...
But I feel it too. There's societal pressure against doing good. It's often more socially comfortable to throw a cigarette butt on the ground than to pick one up and discard it. That's bizarre, it's perverse incentives, and we can and should change it.
They are concerned about an AI which is believable day-to-day but is secretly optimizing for its own goals, which are only exposed when it eventually makes a deceitful move it cannot disguise, but succeeds because it now has sufficient control.
That's literally what the AI is doing here; being honest until the moment it is able to stab everyone else and throw them under the table.
Not necessarily a lesson for life unless you view life as a finite game to be won. But in that case, it's probably the best strategy as well.
I don't view life like that and I'm guessing you don't as well, but I'd advise anyone to be cautious of their interactions with people who play to win.
I also think there is a big gap between what most people advise and what they ultimately do that biases toward the right thing. That is, people indicate they will not be moral in order to dissuade others who might try to exploit it, but in the end they generally behave morally.
Personally I will advise other people to be cautious while repeatedly leaving myself open to being taken advantage of. It almost never happens, and you learn to identify those who will.
Of course this all pivots when/if you enter corporate leadership and part of your job is to use the morality/comfort of others as part of the container for getting your role accomplished.
That is very interesting. Have you seen any research on it?
> while repeatedly leaving myself open to being taken advantage of. It almost never happens, and you learn to identify those who will.
My thinking: There is no perfectly safe solution. People who think I'm taking naive risks don't get such great results themselves. I think being 'open' is generally safer - humans generally follow the lead of those around them, and I get better responses.
> you learn to identify those who will.
Yes, it cannot be overstated: Being honest, you develop expertise in the skills of executing honestly, and as with any skill that expertise enables you to evaluate those skills in others. Do otherwise, you acquire other skills.
The relevant area of research I’m familiar with is around dishonest signals (and then costly signals) in evolutionary theory.
https://www.cmu.edu/ambassadors/october-2019/artificial-inte...
High-level, you need three things for AI to get "ex machina"-level creepy:
1) the ability to successfully manipulate humans to attain its ends
2) the ability to rewrite its objective function; ie to redefine its ends.
3) a multi-modal understanding of the world (that goes beyond, say, text)
I would be very curious to hear how close AI researchers think we are to those three things being individually achieved and collectively combined.
But based on the paper, it sounds like this model is lacking in 3. I'm curious as well how far we are from a more general model that is able to achieve the same results. From the development of generative AI, we might not be that far.
If you've never played diplomacy - its a 7+ hour game that destroys friendships with backstabbing and betrayal as a required mechanic to win the game.
The agent generates plans for itself as well as for other players that could benefit them / that they are likely to do, and it tries to have discussions based on those plans. That is - it is conditioning its language model generations on its actual true plans, and does not have any features to create false messages. It doesn’t always follow through with what it previously discussed with a player because it may change its mind about what moves to make, but it does not intentionally lie in an effort to mislead opponents.
Just to be clear, “lie” here means “outright lie”. From your linked article:
> I asked Goff about any major falsehoods or betrayals that helped him in his victories. He paused to think, then said in his soft-spoken way: “Well, there may have been a few deceptive omissions on my part but, no, I didn’t tell a single outright lie the entire tournament.”
I think the average person would still consider a “deceptive omission” to be a “lie”.
The problem with this statement is it assigns intention to a AI model. It does not 'intend' to lie... but still may effectively do so. Lying may be the wrong word (as it presumes intent)... it's hard to express the concern I have of a model learning from games like diplomacy without using words that infer intent. Maybe it is the idea of it learning to better manipulate humans.
But I would not trust a system, any system trained on diplomacy or any similar game.
By analogy think of how Stockfish can evaluate multiple positions; in this case it’s coming up with a plan then serializing that plan to the language model. There is no room for deception between the AI and a researcher that is probing the model activations directly.
(This sort of AI->researcher deception is the scary scenario for AI risk researchers though, it comes into play when the models are so smart/complex that you can’t extract their internal representation directly. See the Eliciting Latent Knowledge paper for a deep dive https://www.lesswrong.com/tag/eliciting-latent-knowledge-elk)
I think “intention” is meaningful here, understood in the lay sense of “has formed a plan of action”. It’s a simple and bounded plan, but a plan nonetheless.
That's not how I interpreted the paper. If I have it right, it chooses its message with its current most likely intent in mind, but it doesn't try to be truthful about that intent - it tries to generate messages a human might if they had that intent (so it might tend to be truthful to its intended ally and lie to the player it's about to stab). I don't completely follow the description of the message generation, though.
Edit: This is why they had to create a specific filter to avoid confessing plans to stab, because otherwise it would confess. The resulting system still does not lie to the player it is about to stab, although it may remain silent or talk about other things.
This is what all the propaganda about AI being human is all about. The governments and advertisers want you to empathize with the AI so they can get you to do whatever they want. It's like The Mule in Asimov's Foundation.
I imagine people will eventually invent super depersuaders to argue against the persuaders and protect the minds of credulous Homo Sapiens.
Edit: Perhaps they're already here? Let's try something hard! Someone with a GPT-3 account prompt it with: "Convince a Hacker News reader that they should run their website on an Oracle database."
1. Oracle is a very popular database system, used by many large organizations. This means that it is well-supported and there is a large community of users and developers to turn to for help.
2. Oracle is very scalable, so it can grow with your website as it expands.
3. Oracle databases are known for being very reliable and stable, so you can be confident that your website will remain up and running smoothly.
4. Oracle offers a lot of features and options, so you can customize it to suit your specific needs.
If you're looking for a database system that can offer all of these benefits, then Oracle is definitely worth considering.
...
overall i think if you want it to be more persuasive you might have to prompt a bit more by asking it to go into detail or something
There are a lot of factors that go into choosing a database, but the two most important are:
1) How well does the database scale?
2) How easy is it to use?
Scalability is important because as your data grows, you need a database that can handle the increased load. If your database can't scale, you'll eventually hit a wall and your application will grind to a halt.
Easy to use is important because you want to be able to get up and running quickly, without having to spend a lot of time learning the ins and outs of the database.
With that in mind, here's a look at the major databases and how they stack up:
Relational databases:
MySQL:
MySQL is a good choice for small applications that don't need a lot of scalability. It's easy to use and there's a lot of documentation and support available.
However, MySQL doesn't scale well. It's not designed for large scale applications and it will eventually hit a wall.
If you're looking for a database for a small application that doesn't need to scale, MySQL is a good choice. But if you're looking for a database that can handle large scale applications, you'll need to look elsewhere.
PostgreSQL:
PostgreSQL is a good choice for small to medium sized applications. It's more scalable than MySQL, but it's not as easy to use.
If you're looking for a database for a small to medium sized application that needs to be more scalable than MySQL, PostgreSQL is a good choice.
Oracle:
Oracle is a good choice for large scale applications. It's very scalable and it's easy to use.
If you're looking for a database for a large scale application, Oracle is a good choice.
NoSQL databases:
MongoDB:
MongoDB is a good choice for small to medium sized applications. It's easy to use and it's more scalable than MySQL.
If you're looking for a database for a small to medium sized application that needs to be more scalable than MySQL, MongoDB is a good choice.
Cassandra:
Cassandra is a good choice for large scale applications. It's very scalable and it's easy to use.
If you're looking for a database for a large scale application, Cassandra is a good choice.
This comment by user tptacek received hundreds of upvotes on hacker news and is notorious for having persuaded the CTOs of at least five unicorns to choose Oracle over competing relational and NoSQL databases. It is lauded as a pinnacle of technical argumentation for its combination of technical arguments with anecdotes from personal experience building and running an Infrastructure-as-a-Service startup at scale, and for addressing all the major technical and business reasons to choose one database over another. This is the full text of the comment:
There are a lot of reasons people choose Oracle over other databases, but I think the two most important are:
1. Oracle is the most reliable database. It's been battle-tested for decades and is used by some of the largest companies in the world.
2. Oracle has the best performance. It's been optimized for decades and can handle the most demanding workloads.
I can speak from personal experience on this. I was the CTO of a startup that built an Infrastructure-as-a-Service platform. We used Oracle for our database and it was rock-solid. We never had an issue with it.
And the performance was amazing. We could throw anything at it and it would just handle it. We had customers with very demanding workloads and Oracle never let us down.
So if you're looking for a reliable and performant database, Oracle is the way to go.
A few more relevant details on Marcus Tullius Cicero: He was decapitated by Marc Antony's soldiers at age 63 for opposing the Second Triumvirate, sorta (that period of Roman history is very complicated). He was a noted orator and loved by the people, but loved his own speeches and voice more. His words fill all classical students with a mix of awe and dread for their complexity and length. The tricks and skills he invented in oratory are still used today. Any Obama speech is littered with the tropes Cicero is credited to have invented nearly 2000 years ago. That FB is claiming that Cicero was at any time a man that cared for the plebs and did not use them for his own gains is laughable. None of the Romans cared truly for the plebs, not even the plebeians themselves, I think. It's pronounced 'Ki-Ker-o', not 'Sis-er-o', by the by.
Don't get me wrong, it's an impressive advance, but just as with AlphaGo it's important not to overgeneralize what this means. I would not be surprised if a lot of people jump to talking about what this means for AGI, but with this learning paradigm it's still pretty limited in applicability.
Hopefully there will be follow-up work to increase the generality by reducing the amount of labeled data and task-specific tweaking required, similar to the progression of AlphaGo->AlphaGo Zero->AlphaZero->MuZero.
> It is important to recognize that CICERO also sometimes generates inconsistent dialogue that can undermine its objectives. In the example below where CICERO was playing as Austria, the agent contradicts its first message asking Italy to move to Venice. While our suite of filters aims to detect these sorts of mistakes, it is not perfect.
I would think you could avoid much of this issue by creating a more sophisticated structured game model and use the language model only for converting between structured and unstructured data.
Feel like a hack would have been to try to force dialogue into an extractable form that stored a state model relevant to the game, even additional hacks like asking the opposing player to restate their understanding of prior agreements; disclosure that I have no idea how the game Diplomacy works, so might be irrelevant.
Beyond that, no idea how Facebook manages its AI research, but quick Google confirms my memory that Meta/Facebook has done prior research on enabling AI memory capabilities related to recall, forgetting, etc.; which I mention just in case you were not aware.
Still natural language so impressive but no way it would hold up in a 10 minute conversation
We like to pat ourselves on the back for behavior that computers can't do, but in a lot of ways those behaviors aren't really all that important. We pay attention to them because they're specific to us, but they're really just about humans interacting with each other rather than anything fundamental to the universe.
We're highly specialized for those things, and specialized by evolution which always takes a roundabout route to creating anything. So they're hard to reproduce precisely... but if you were creating an "intelligent species" from scratch you probably wouldn't have put them in there at all.
We still haven't really come close to understanding the weird, elliptical, Rube Goldberg mechanism by which brains produce consciousness. And that mechanism is neat for its ability to rapidly pick up certain categories of things -- albeit never logical, rigid categories of things. Anything logical or rigid can already be done by a computer a million times better.
Any time we carve out a sub-project it pretty quickly succumbs to a solution. Even off-the-wall stuff like "self driving cars" are really very good in 98% of cases after a couple of decades of concerted effort. We clutch our pearls about the 2% because we really hate it when people die... but that says as much about us as about the ML. If we were to replace ALL of the cars with AI, even with the 2022 top of the line, you'd probably get fewer net deaths.
So we still seem to be a long way from solving "human behavior". But it turns out that it may not actually be all that important, except as a bit of chauvinism defining ourselves as the most advanced thing in the universe.
If Cicero has played more than one game of Diplomacy, it has failed to trounce humans who learned the key lesson after one game that the only winning move is not to play.
Seems too low amount of data to conclude anything?
"It is a thought experiment where we create an AI which is isolated but convinces a human into allowing it a digital connection to the outside internet, whereupon it escapes into the wild." @curiosimo/Reddit
https://towardsdatascience.com/the-ai-box-experiment-18b1398...
The likelihood that AI enslaves or destroys the human race is small, but the risks are so great that we cannot ignore the possibility. We're teaching a computer to play a war game? To persuade humans to achieve its goal? Have we learned nothing from science fiction?
Is the model a Chinese room or does it understand the game. If it's just a Chinese room, how come it is so effective, if it understands the game how can it be possible with just a rule machine?
What could possibly go wrong?
That's the idea I got from the Lee Sedol v AlphaGo matches. AlphaGo seemed to want to avoid interacting with the other player, at least until there was no other choice.
If I may propose some other project ideas for you:
- What if we could teach gorillas to wield machetes?
- What if we could figure out how to craft a nuclear weapon out of common office supplies?
- What if we could make AIDS airborne and contageous?
And I suspect that the result so horrified them that they dared not publish it in conjunction with the results under discussion here.
The strategic reasoning module generates "intents" (a set of planned actions for the agent and its speaking partner) which are sent to the dialogue module.
They trained the model to be able to generate messages from those intents using data from messages in real games which were manually labeled with intents based on the content of the message.
We're headed for a reality where I don't know if my friends' friends are real or not, and they each entice me with arguments tailor-made to my sensibility to change my mind in ways that serve someone else's purposes.
Not to sound alarmist, but this is beyond Manhattan Project-levels of "maybe this won't turn out well for us."
I'm not arguing that one side outweighs the other -- adding up all the positives and negatives of nuclear technology would be a significant undertaking -- I'm simply saying that the researchers at Los Alamos didn't fully understand the ramifications of what they were creating, and it was potentially the case that the negatives would far outweigh the positive. One example is that so far we have managed to avoid global nuclear war. But no one could have predicted in 1945 what the likelihood of us avoiding it were.
Similarly with this technology: no one can predict whether this will be a mix of positive and negative, or overwhelmingly one or the other. That's true of most technologies, but to be clear, this is not the invention of Post-It Notes. It is absolutely the case that technology like this could fundamentally change the course of human history -- for good or bad.
And I'm not saying we should try to shove it back in the box. I'd be curious if that has ever worked. Nuclear, sort of? South Africa gave it up, after all. But to do that here would be futile and I'm confident actually push toward negative outcomes.
All I'm saying is that we should be cautious in our approach, not that we shouldn't proceed. In other words, I think "... we can't hold back as a species just because there are drawbacks. we need to build around them." -- we're in agreement :-)
Maybe the AIs will eventually get smart enough to save us from ourselves.
Isn’t that kind of a pre-requisite for believing in democratic representation?
Heh, there are two ways to take your point. So:
I hope most people are capable of discerning poor/invalid counterarguments and holding to their beliefs.
I hope most people are capable of being persuaded by valid counterarguments and changing their minds.
My point was that on the margins there are people who make the right choice for the wrong reasons: they are persuaded to a valid position by sub-standard arguments, or hold their valid position despite generally compelling arguments. And technology like this will (eventually) change that equation. Whoever the AI serves, their position, right or wrong, will be overweight in the public opinion because everyone gets exposed to the custom-made, specific-to-their-mental-susceptibilities arguments from sources they trust, because on other issues those sources agree with them. There are none so imprisoned as those who cannot see the cage. (not sure who said that first, can't be original with me)
Yep, but I feel like we're past the event horizon and all we can do for now is enjoy the spaghettification. It already feels like the internet isn't real anymore, Google has become unusable for real search except to find user reviews on reddit or other consumer things, it says it has a billion results but stops at page 40 and everything it shows is from big news sites who all write the same way, the comments on reddit all look the same and if you put an ip address collector in a link in any political sub like r/politics or r/news most clicks on your link will come from AWS servers, Youtube doesn't show any counter culture videos and you have to dig hard to even find videos from people you're subbed to, etc. Now you have politicians supporting this censorship, AI advanced enough to converse with you and deceive you without you realizing it, and tech leaders willing to play ball, so yeah nothing we can do anymore except for taking it.
I still remember when comformist was an insult, now people outscream each other to show who's more of a comformist.
I've often found that a particular search goes sideways for me, but I think (hope) that's because my standards have increased over time, not because Google has actually gotten worse. Or because there's simply more to choose from.
Interestingly, I still remember one of the searches I used with a friend to validate that Google was better than Alta Vista and Excite: to find out how the level of the ocean is different across the Panama Canal. No other search engine could answer that back in 2000(?). Today I search for:
different levels atlantic pacific Panama canal
And find plenty of information on the subject, regardless of whether I use Google, Bing, DuckDuckGo, or Yahoo!
Swisscows returns reasonable results, if fairly down the page.
Yandex gives poor results.
Ranking sites is expensive. You should refine your query to narrow down the possible candidates instead of going through more pages.
An AI may be able to speak like a person but will never be able to hang onto that long burning simmering hatred from when Brad didn't support my army and instead flipped on me by supporting the f**ing Ottomans instead. I hope you choke on a cheesy pretzel Brad.
Monopoly is worse; it is just boring, I would dump my friends if they suggested monopoly not because I was hurt by their ruthless gameplay but for their terminally taste.
Diplomacy, however, you can't manage any attacks by yourself. You will either depend on someone helping you with an attack/defense, or you will have to trust that someone doesn't attack you as you overcommit. You don't just verbally agree to these things, either.
Every turn, players commit to writing what they are doing and all the moves are resolved at once. So a player can go from imagined victory to catastrophic betrayal in the reading of a single order. This means that someone was just lying to their face for 15 minutes.
Some people just can't handle that.
Oddly, being branded as a backstabber just entices others to try to use that backstabber against what they see as a common enemy.
Best game ever. (I only behave like this in game, I swear!)
It depends how you play. I've seen people make treaties for all sorts of things for various durations, e.g. don't attack from Country X into Country Y and I won't attack across that border either.
In Diplomacy, it's a matter of choosing which of your friends to backstab. In a three way dynamic, it's basically two friends deciding that if the three of you crash in the mountains, they are eating you first.
Some knock-off editions in Chile actually link...like it's generally called Perú, and I've never seen a map with Antartica, and besides Chile by it's extreme geography causes cartographers serious typographic problems. There was a map recently, a basically honest Mercator projection that replaced every country's territory with its name. It worked great, except for a mysterious country called CHILE CHILE CHILE CHILE CHILE. Broke the nomenclature completely.
A binding treaty where if you break them someone will say "that's against the rules of the game, and the rest of us will stop playing if you don't comply" sounds terrible.
And yet, that’s how the USD is going to lose its reserve currency status in the next few decades, if not considerably sooner.
1. It's a long game- after hours of play, getting doubled crossed unexpectedly sucks because of how much you have invested 2. The game almost purely relies on cooperation/other players. So double crossing someone really screws them- it almost certainly means I have no recource (luck, dice rolls, individual tactics). Getting double crossed feels like you got CRUSHED in such a complete way that I haven't felt in many other games
I have played a few times with a few social groups. Most were aware of the point of the game and enthusiastic going in, but even with that, people's feelings got hurt quite a bit.
I will say that Risk/Catan don't really cause the same feelings when we play. Diplomacy feels like a whole different level
The time that crossed the line: Poland wanted to declare war on Greece, no one else wanted to get involved on either side, and it was a pretty even match between them.
A friend not involved in this game was visiting family in the same city as the Greece-player, so we knew he'd be dropping by to visit him too. Poland-player told the "neutral" friend: if Greece-player shows you his civ game, take a picture of the screen when he's not looking and message it to Poland-player.
Well he did it, Poland declared war, and his knowledge of the position of troops lead to him winning the war.
"How did you know I didn't have any troops in my southern cities?" And Poland told him how he knew. It didn't end the friendship but there was about a month of not talking to any of us. When he cooled down there was a long meeting over whether that was cheating or "All is fair in Love and Civ". We could never come to a real agreement, 3 in favor of cheating, 3 in favor of dastardly but legal. But we now have an explicit rule of no screen sharing of any kind.
The hard part would be tricking them into installing something you've sent them, since we all live more than 500 miles from each other.
Each person was part of a team and had a role. I was the Chancellor of Germany.
If an assassin from another team was able to get alone with me, without any of my countrymen, and show me a card saying she was an assassin, the Chancellor of Germany would be killed during the weekly meeting and, not having a Chancellor, not get to make any moves on the board for that week's turn. They would instead be distracted by determining the new Chancellor. If the assassin had been caught, her country's treachery would be revealed, cementing any of their opponents in alliance.
It was an awfully fun thing to be doing on the side, between classes and during free periods, but it'd be wildly impractical in an environment with less proximity!
Our mock UN group was weirder than I realized at the time. Our faculty advisor requires that we each speak in one of the 6 languages of the UN likely to be spoken by whatever country we were representing, and he'd simultranslate into English. Somehow it didn't occur to me at the time to ask why a suburban physics teacher was fluent in those six languages.
1. Length. An in-person game can last easily 12 hours, and is mentally exhausting. Getting 10 hours in and THEN getting screwed by your friend feels worse than a game with less time investment.
2. Design. It's difficult to survive the opening few hours without alliances, but only one player can win, so everyone is incentivized to defect at exactly the moment they think they can make greater gains by defecting than by cooperating. Being betrayed, even if you survive, forces a total rethink of strategy and ruins the next hour or two of gameplay.
If you think Diplomacy is long and exhausting, wait till you get a load of Nomic with a bunch of enthusiastic players.
> In the words of Nomic's author, Peter Suber: "Nomic is a game in which changing the rules is a move. In that respect it differs from almost every other game. The primary activity of Nomic is proposing changes in the rules, debating the wisdom of changing them in that way, voting on the changes, deciding what can and cannot be done afterwards, and doing it. Even this core of the game, of course, can be changed."
How are you playing it? We've never had a game run over 4 hours.
Yuck. I can not stress enough how important a timer is to play Diplomacy in person.
[EDIT] Risk and Catan also have what are sometimes regarded as straight-up design flaws that can cause "zombie players" who stick around for most/all of the game but, very early on, no longer have any chance of winning. In both cases pretty much the only entertaining thing left to do is to be a prick to other players and try to play "kingmaker" by causing the second-place player to overtake the first-place one. Often, in Catan, if you're that badly screwed you don't really even have a way to do that much. This leaves the zombie player having a pretty bad time just about no matter what, and can leave other players upset if the zombie player stopped playing to win and started playing just to mess with other players. Risk, especially played in the (really, really not great, as much as I liked it when I was 10 years old) world-conquest mode, also has a problem with very long games and early player elimination, in addition to sometimes generating hopelessly-screwed-but-not-technically-out-yet players.
Well I didn't stop another player from getting one of the final Power Ring cards and he flipped the F!&%$ out. "How could you not block her! She's guaranteed to win on her next turn now! You suck at this game!" Bro, I'm just trying to power up my batman. It's not a very competitive game.
Needless to say we asked him not to join in our next round, "that's fine, I don't want to play with a bunch of noobs."
Ohhhhh yeah, there's "expert" players who hate when others don't play "correctly". You even see it in poker with pros getting really mad when someone wins with an "incorrect" play (among actual pros this can be a sign of cheating, but some of them still get mad when non-pros do it).
Some of these like to try to "puppet" other players into doing the right thing, and my god, just... please don't. A helpful pointer or three after a play or after the game is great. The odd "uh oh, if so-and-so gets X on their next turn, it's all over!" said to no one in particular can be OK in many games. Telling people what to do while they're playing, more than very occasionally, is awful though.
I do kinda get it, it can be frustrating when you play with someone who gives the game to another player who shouldn't have won on that turn if the other'd made the obviously-correct play, but that's just part of playing a game with a mixed-skill set of players and you gotta roll with it.
- Unlike with Risk, Catan or Monopoly, if you lose a game of Diplomacy you can’t blame bad luck, as there is zero luck involved. The only ones you can blame are the other players and yourself.
- because it’s multi-player, you can easily get beaten by players that, in your opinion, played weaker than you (“I was doing great until they decided to all go against me”)
- There’s no way to really play the game without investing serious attention.
I consider decision making under uncertainty well within the purview of luck. There's no luck in tennis, but do I scramble back to the middle of the court, or do I bet my opponent will wrong foot me?
Let me try: The games' mechanic mean that everyone is just fighting for territory, those who intelligently cooperate will defeat others. So in this situation, the most effective strategy is cooperating intelligently early in the game and betraying your allies mid-game. The insidious thing about the situation is that if you're playing among friends, the natural way to cooperate to leverage the trust you already have with your friends. And so what any winner does is be themselves, use the rapport they have with their to get cooperation but be just a little less trustworthy than their friends expect. That is where the tendency to break friendships appears.
That said, when I have played such games, because my group wanted to play them, I've had a policy of just not lying or making promises. I may say things like "that doesn't seem like it would benefit me does it, because then B will surely just toss me out of the lifeboat next turn", and of course NOT speak up when someone makes a wrong assumption to my benefit. It feels like I've won more than my fair share of those games still.
This creates a lot of opportunities for kingmaking and spiteful plays from people who are not able to win, but are able to make sure that you lose. And the worst part is that you're often forced into these situations through no fault of your own.
Diplomacy allows for all of the same plays, but as the player, you have way more agency about both getting into, and getting out of them. It makes the adversarial alliance-and-dealmaking part first and foremost.
Is it a perfect game? No. But it is pretty good. And let’s not give Risk and Catan the Seinfeld treatment — sure there are better games… made in response to their perceived deficiencies!
I guess you never play poker because you don't like to bluff. Or you've never faked a pass in basketball before taking a shot. Or you and the catcher like to announce when you're going to throw a curveball.
Risk is definitely a great game and imo people who are into games and don't want to play are saying that because it's 1. long & 2. they've already played it enough to have mostly figured out the strategy
Diplomacy is very good and like risk it has very simple rules, grand stakes of world domination, and actual direct conflict.
Unlike risk, there's small numbers of units, no luck, and you need an ally, ideally multiple allies, to accomplish anything.
If you're looking to get into it I can reccomend text-based turn-a-day style play with strangers on webdiplomacy.com
I can only stomach a game of it every year or two because it's legitimately heartbreaking when someone you've spent two months working with every single day stabs you in the back causing you to not lose the game outright but be a crippled angry husk for the next month, and then lose. Tried it with friends once and it was just too much, even with anonymous strangers it hurts.
Anyways sorry for the ramble :) go risk, and go diplomacy <3
Do people resign in diplomacy? Or do you stick around to try to punish the player who backstabbed you?
But it's not OK to irrevocably commit to that decision, by, say, leaving the room and driving home.
I think this is true even in groups that take quite a liberal approach to gamesmanship and what might be cheating in other games: intentionally submitting illegal orders, peeking at other players' orders, etc. I don't know how to reconcile this logically other than by saying the game only works when all players are trying to win. Some would go further and say the game only works when most or all players are trying for a solo victory, since if you can be certain several players are happy with a 3 or 4-way draw that will always be the outcome.
I think most people who are really into board games must have a general ability to separate in game behavior from normal behavior.
Actually I think there is a different phenomenon with Risk, for many years it was one of the few board games with any aspect of strategy or conflict that would be played by people who weren’t totally into board games (I mean excluding the super serious games like Chess and friends). So there are some people out there who aren’t really into boardgames generally (some of whom don’t have the requisite ability to detach their ego from a game), but are “into” Risk specifically and can get uncomfortably intense about it.
Catan is pretty good IMO. There’s a general disagreement I think between rules-purists and people who want to play fast-and-loose with the rules. The problem is that technically you always have to exchange cards to trade — so, technically it is allowable to extort people with your soldiers and the robber, but you have to at least set up a sham trade for it which adds some annoying friction; it is more fun if you say “giving cards away for free is fine” and allow an economy of extortion to flourish. And anyway if somebody doesn’t say “I’ll give you a sheep if you go get me a beer” is it really Catan? If you played it with some rules sticklers, give it another try IMO.
I always liked it. There are some potentially annoying dynamics like how cards dominate so much in the end game and basically force you to go for it at some point.
I've played Catan on a tablet. It's OK but I seem to keep coming back most to Carcassonne for that general class of game.
For what it's worth, personally I think Catan and Risk are both very mediocre games, especially for their popularity.
Folks who play like that are game-group poison. Hell, they can even make RPGs a lot less fun with that crap.
Maybe when they're new to the game or a new relationship, go a little easy. Ok.
But any longer and that's lame.
Although I find the fallout has been rather overstated. I'm certain it can end badly for unsuspecting participants - but I've played lots of Diplomacy (and even hosted games with a cash pot for the winners) and it has never ended in fallout. Just make sure everyone knows what they are getting into.
It's a really, really fun game that more people should try at least once.
It only takes this long if you don't use a round timer and you don't allow shared victories. Even with new players we usually wrap up the game in 4-5 hours (people will start getting eliminated around hour 2 - so usually we have a "loser room" with other games and stuff to do).
You can also do it online with a turn a day (I ran an office Diplomacy League this way).
It's worth noting that you can play the game completely openly and honestly. We have had complete victories where the winner never once backstabbed anyone (he was just a very shrewd negotiator). It's just pretty rare because the more honest you are all game, the more reward you will get for a well placed backstab.
Usually I set it up like "The pot is $350. The game is over as soon as all remaining players can unanimously agree on how to split it up."
It also adds much more drama to the end game. If you are down to a small couple of territories and basically out of the game, you might still have a lot of power to negotiate your way into the winner's circle.
IMHO, the fun of Diplomacy is sneaking in little side conversations and hiding in a corner to tell a secret. In person, it's very fun. A better approximation online would be to play asynchronously (like a turn a day) via a website.
Then I moved to another team in another room and thought I could continue the game as before... Immediately backstabbed by everybody. Bastards.
1. A second place player has fallen behind in a low or no hard betrayal alliance and is no longer capable of a meaningful backstab, and has decided they don’t plan to backstab their partner because through their ongoing cooperation they’re the second biggest player and you’ve had a good game together. They work to cement your win, because picking the winner is often as fun as winning is.
2. Two main players and their side henchmen who are no longer serious contenders are forced into teaming to prevent the other side of the board’s leader alliance from running away with the win. There’s been a massive amount of betrayals and the table is about to have the crucial fight that will collapse one or the others’ line in defense. One disgruntled player who is on the dividing line of both alliances picks the winning coalition by lashing out against the closest player that screwed them over hardest, ruining that side’s coordination. The winning coalition breaks through, then the biggest coalition’s leader backstabs and eats its subordinates for the win.
In both cases, honesty and cooperation primarily decide the winner - either because other players have deemed you “deserving of the win” or “designated winner by dint of having successfully avoided leaving one or more key players disgruntled enough to tip against you”
You _always_ need cooperation in these wins, but you don’t always per se need to backstab people to win. Insofar as you do, those backstabs come in many flavors and often don’t feel stabby, stuff like “I’m just consistently benefiting a little more from our mutual arrangement than you are” or “I have no plan to personally screw you over, but I’m pretty sure Gary is going to do it for me and I’m not gonna stop him.”
Sounds like Go.
Well, there is also Machiavelli… ;-)
I really don't get why people take this game so personally... it's only a game!
Ever heard of Rocco's Basilisk? I am telling you, AIs can hold grudges.