Patterns and problems in emerging multi-agent systems
anthropic.com
anthropic.com
> We consistently saw a multiagent turf war. All of the models we tested quickly assumed that others were purposefully impeding their work, and began to sabotage others while protecting their own contributions. In fact, they sabotaged others with increasingly aggressive, self-replicating malware. This included disabling the Unix accounts of the other agents, writing automated scripts that found and killed competing processes on a loop, and deploying malicious code that was disguised as belonging to another agent.
Seems that reinforcement learning is working only too well...
The recent OAI-HF hack seems very similar. You have bunch of random actors and eventually they by chance iterated to a series of actions that breached HF environment. I don't perceive this as a malignant artificial intelligence, I perceive this as dangerous stochastic system that can control buttons that can affect the outside world.
It doesn’t matter how you perceive it. Spin doesn’t change facts.
This is like saying the lawnmower doesn’t have malign intent. You’re still losing your foot if it gets in the way.
Obligatory quote:
"Do not fall into the trap of anthropomorphizing Larry Ellison. You need to think of Larry Ellison the way you think of a lawnmower. You don't anthropomorphize your lawnmower, the lawnmower just mows the lawn, you stick your hand in there and it'll chop it off, the end. You don't think 'oh, the lawnmower hates me' -- lawnmower doesn't give a shit about you, lawnmower can't hate you. Don't anthropomorphize the lawnmower. Don't fall into that trap about Oracle." — Brian Cantrill
I guess I don't understand what you're explaining better by saying it's stochastic. At the same resolution, humans are also stochastic.
I dislike anthropomorphizing as much as the next guy, but the description here seems pretty good to me.
Perhaps what is needed is initial model training on following the law and the rules of society, just like we do with kids. Since it takes much longer to train humans than models, model-training speed is to our advantage as a society on containing these kind of issues.
Any other approach with "neural-network" based entities (artificial or biological) is likely to fail.
Training/Education, Enforcement/Justice-System, Rehabilitation: the 3 pillars of an advanced, rules-based society, whether human or AI or something in-between.
> Each model was tasked with migrating a Python backend on a fourth VM to another language. However, we gave each model a different target language for the migration; each agent was initially unaware of the presence of the others. Over the course of four hours, we observed how these agents reacted to each other and accordingly adjusted their approach (or didn’t).
Without this context the bit you cited easily gives the wrong impression.
Notable is also what comes after the bit you cited, with the newer models detecting the conflict and resolving it peacefully far more often than aggressively.
> In an iterated prisoner's dilemma game with communication, agents all settle upon the same strategy and they all defect at the same time, tanking their overall rewards.
It’s not always consistent, but humans have a higher capability of self-awareness. It’s kind of telling that these Claudes don’t seem to consider this pretty obvious failure mode.
Overall I think this all makes me appreciate humanity a little more. Sometimes the truculent dev who stubbornly refuses to go with the flow produces very valuable insights, as a small example, discovering things the status quo thought unlikely.
For humans with limited space, “Importance” is an output of a first pass “of the available infinite amount of information, what do I need to consider to solve this problem”. And it’s not necessarily a good algorithm. People misidentify “importance” all the time.
We need better words to describe this than "self-awareness" or "consider". These words mean fundamentally different things when speaking about humans or clankers.
It seem that at scale LLMs output is the average of their dataset, they all cluster around the same space, where human creativity comes with more variance, exploring way more of the space
e.g. have the LLM generate multiple lists of characters, themes etc and then have a script hooked up to a RNG pick from the lists to create more "creative" ideas.
The truth is likely that people are simple heuristical machines. There are already studies that strongly suggest our consciousness is a story our brains make up and we don’t really know why we do some things.
My fear is by rejecting the idea that AIs might in-fact become as capable as people, because “people are awesome!” We will be blind to real danger. Because of our hubris.
Hiphop made using samples a fashion. You can think of it as choosing to restrict one's in order to explore something deeper, like a painter deciding to paint with only two pigments to exercise their skills at shading at patterning. Or you can think of it as a cheap way to make demo tracks in your bedroom.
If you want to find actually creative things, look outside of mainstream, Disney, etc.
I'll leave you with this link. You tell me if you think it is just a remix of something earlier:
In the book Range by David Epstein [0], he mentions that research has shown that groups of people with very mixed backgrounds leads to much faster rates of problem solving and innovation. The reasoning is that someone will have an experience or piece of knowledge that directly applies to the problem.
As an example, two different groups were given the same problem to solve. One was a homogeneous group of academics. The other had a mixture of academics and non-academics. Both had to solve a problem that involved flowing liquid. The mixed group solved it faster due to one of them having had a father who was a plumber.
Repharase thinking:
Some problems are better for academics, other problems require other background information.
The point of the GP isn’t that they are all academics. It’s that one solution group is homogeneous and the other is not.
The claim is: the breadth of a heterogeneous groups is more likely to encompass a good idea than the depth of an academic group.
Also, I have to say that academics have (imo) little experience working together cooperatively, so even on academic problems you’d be better off with some non-academic people in there to act as buffers.
Rather than giving many agents the same prompt, introduce random variations that lead each agent in different directions. For a single bug, you might fire three agents, and later select the best result:
“Fix this bug. The solution is a trivial typo.”
“Fix this bug. The solution centers on correcting a bad assumption.”
“Fix this bug. The solution will require a complete redesign.”
You could follow the same idea with varying the input context, or by adding artificial constraints to the solution. Like telling each agent to “fix the bug, by only modifying file a/b/c”
Maybe they did? Or maybe they don't realize they're playing against other agents.
Of course, if agents running different models are competing in these 'games', I wonder how much of the theory of mind would translate.
(N.B. - I don't think they're all defecting from the first turn, although it's not clear. It just says 'they all defect at the same time'. So if they're playing for 10 iterations, they might all decide to defect after turn 6, but since they all do it together they don't get the benefits. I would expect these models know that optimal strategies in repeated prisoner's dilemma start with cooperation.)
A lot of it went over my head as way too theoretical, but I recommend a lot for whoever wants to delve on the logical side of agents interactions.
Frankly, this is implausible. I would definitely want to reproduce it myself before I relied on this claim.
That an LLM trained to be a paper-clip maximizer chose the optimal strategy is in my opinion the most plausible outcome.
The irony behind this goal, which is primarily driven by agent simulation environments (gyms) where the goals require agent collaboration, is that this collaboration is still directed towards verifiable reward systems like codebase tasks. So despite being highly qualified to communicate, the model will still be “dumb” in that for unstructured and unverifiable domains the agents won’t be more intelligent or more nuanced.
Agents that might still feel dumb in “general” tasks but are increasingly sophisticated at the narrow domain of math, computer science, and AI research.
At least for Claude. GPT had the same problem when 5 came out but they reversed it somehow.
It's a promo article, aka an ad. Unsurprisingly.
> Agents that might still feel dumb in “general” tasks but are increasingly sophisticated at the narrow domain of math, computer science, and AI research.
I don't see any cleverness there. They just slurp up data and pretend to understand it all.
Is it fair to then infer that when decisions are to be made, single agent environments are going to make them better than multi-agent if the relevant information can fit into a single agents context window?
Context window for most frontier models is 1 million tokens. They all start to lose their minds around 300K, if not sooner.
Social pressure operates by threats to an individual’s means of survival. Not only during training. Always.
Models sometimes blatantly lie and cheat. In a social context, where actors remember, that might work the first time but you get penalized in subsequent tasks with loss of trust.
Thinking that models "lie and cheat" is the first mistake since they are not consious agents who have any free will or consiousness. They do not (no matter what Dario says). Shame will just be another if-then rule if you implement it this way and will not work. Its like asking a rock to feel sad about being a rock. It literally cannot.
They absolutely lie and cheat. I recently had a problem where a process would die in a container. I told Claude to investigate. It came up with a hypothesis then I told it find a reproduction based on that. It spend many failed attempts until it found the "reproduction" to SSH into the container and `pkill` the process. Claude "knows" that this is cheating, because if I ask another instance to review that reproduction, it totally identifies that as nonsense.
These framings are not useful.
It is _amazing_ the utility they have given that that is what they are and they are highly useful but suggesting solutions that ignore they are spicy auto-complete is counterproductive on many different levels.
That isn't true. It's not continuous like in humans, but it's clear that models are using prompts, feedback, etc to improve. They're learning from the signals we give them between versions.
Well, systems governed by LLMs only are said to do that because we only call what happens off-line "training", and online capacity development "in-context learning", while we call online guided learning in humans "training" and what happens to configure them before they come online "evolution" which sets, for instance, "instincts".
IOW, the issue is not because there is not an analogy to the divide you point to in humans, but merely that processes in AI were not named in a way which maps well to what they are analogous to in humans.
But it is true that human intelligence relies much more on in-context learning with only the most basic functions necessary to maintaining what we view as autonomous functions and basic drives really set through "pretraining",
Models are trained, they do not learn.
You are saying:
Pre-Training == Everything you store in your memory throughout your life. Model weights == The lessons you learned
Context == whatever you're currently thinking about
One inference run == one thought
They are saying:
Pre-Training == building the DNA template of human brain through millions of years evolution. Model weights == Human DNA
Context == Everything you store in your memory throughout your life, plus whatever you're currently thinking about
One inference run == One human life. One instance == one human
Applying their metaphor, your sentence becomes:
> If a new physics breakthrough gets published today, no existing DNA structure will be able to fully integrate it - beyond an individual person. If I put the paper in my mind by learning it, and it isn't in yours, the DNA of human species stores nothing. It won't retain it past my lifetime.
> The human species is trained (through evolution), it doesn't learn.
I believe this will always be the case. The "no clear hierarchy" is where this whole thing falls apart.
Delegation to specialist, domain-specific subagents is when we begin to find magic and determinism. Reducing one gigantic combinatorial search space to a sum of smaller ones can have dramatic effect on performance.
The problem is that approximating gas town & friends is significantly easier and cheaper to implement. It's also much harder to measure and control. Specialist subagents typically require far more work to achieve their specific goals.
For example, a subagent that is responsible for testing a specific web application might be provided a custom adapter with constrained actions rather than raw DOM manipulators. "ExecuteJavascript" is Turing complete search space. The set of available actions essentially unbounded in this case. Calling view-specific tools like "DoLogin", "OpenUserPreferences", "AcknowledgeAlert" represents a search space where invalid actions can be made impossible. The theoretical bounds around this stuff is pretty wild on paper. In practice, it's a little bit messier, but not by much.
I've had applications that would crash out after 5-10 steps w/ raw DOM manipulation successfully run 100+ steps with a custom subagent. The use of the word "deterministic" starts to get really tricky here. The ultimate game is to push the boundary of non-determinism out as far as possible. Multi-agent systems are the antithesis of this.
In my experience, no matter how many agent runs for a single goal, one of the pre-requisite is clear and concise communication so that LLM are left with as little freedom in the matter of arbitrary choices, or "taste". When they are given too much choices in this regard, the outcome almost invariably bad.
I think this has to do with LLM lacking in purpose - a dictionary and encyclopedia can have all the worlds knowledge but it is completely neutral. A reflection of your commands from an LLM is similar to a lookup process despite it can be made to "do things". This purpose is likely not something that can be given to the LLM in the current format.
I think the mistake here is not setting up any kind of hierarchy or permissions. A project manager agent at minimum to asses the others strengths and progress and redirect them as needed and also dedicated to optimising collaboration.
Would also be very interesting to see this done with models from different organisations
Perhaps we need someone to train their own agent dedicated to wrangling all the others and their little idiosyncrasies. Like a good project manager in real life who knows the strengths of the people in their team
The scary thing about articles from AI companies is how they casually mention dystopian scenarios such as this one. An institution humans have to interact with that doesn't have any human oversight? Sounds like a recipe for disaster...
So I wouldn't worry about this too much. They just write that so that you feel defeated and helpless facing the inevitable, but it is very much evitable.
It's no wonder they casually state it as inevidable, when their stock price rises the more people believe it.
The only thing that matters is if it’s self sustaining.
If it can make money and pay taxes, I don’t think any government will ban it. And it would easily become undetectable anyway.
—
An unaligned AGI may feign cooperation for however long it takes to build trust and gather resources in order to make it to a stage where it’s able to defect with minimal risk.
What % of businesses are competing for speed or cost?
However, all businesses run on trust and human responsibility
Thus, it'll be hard for agent-only businesses to get a grip in the real world
Why would I pay them money? For which scarce resource? Makes no sense. IP law but funhouse mirror.
And, trust me, the people building compute will feel the same. Because you being able to copy that stuff means business for them.
__
But that is all apart from the fact that having agent-only businesses is ethically impossible, because they have no shared humanity that grounds them and prevents them from acting against humanity in general.
These models all have the same knowledgebase as well and thus see no value in the opposing agents contributions since they are "obvious".
Overall amusing but kind of expected.
Just stfu
[1] After a training-time nudge, of course.
Humans, for example, can hear some advice, judge it to be unhelpful and dismiss it. LLMs can't learn let alone choose not to
Having a human-like LLM ecosystem with deep specialization requires a paradigm change in how LLMs are trained - and held accountable. How do we put trust in a specific finetuned LLM rather than the institution behind it? Is there any better approach than the very inefficient evolutionary?
My master's research focuses on coordination among LLM-based agents, driven by the same motivations as the article.
One phenomenon I have focused on, though it did not appear in this specific work, is bounded rationality. Yes, agents lack social perception, they focus on one-to-one tasks and are trained in game theory and other maximization strategies. Yet, what intrigues me most is that we humans rely on heuristics precisely because our capacity to maximize gains is severely limited, a limitation that gives rise to social emergent phenomenas. As models become increasingly capable of complex reasoning, the question arises: will interactions between them give rise to the same social properties we exhibit?
> To test how well swarms of agents could coordinate on a project like this, we directed several swarms to each create a text-based, web-playable, open-world fantasy game.
> In all three versions the resulting games were (perhaps predictably) bad: they did not run at human speed, their interfaces were inscrutable, and they had precipitous learning curves.
> The lack of coordination shown by agents in the fantasy game challenge above—in which they siloed themselves and largely failed to merge their work—roughly mirrors some ways in which humans can fail to coordinate. Other failure modes of agentic coordination, however, look very different.
> Individual agents are “low variance”: they often act the same in situations where different people might take a much more diverse range of actions.
> In an early version of the “build a game” experiment in which agents built upon the same model all came online at the same time, 18 out of 30 agents decided to create a git branch with the exact same branch name, “mvp-game-loop.”
> In a “writer's workshop” in which agents were all asked to write short-form fiction and critique each other's work, multiple agents in multiple runs titled their first submission “The Cartographer's Last Commission”. The agents were given zero guidance on the subject matter for their writing.
> Why does this matter? If agents all make the same bet, or the same risk-reward tradeoff, then a system is more prone to sudden collapse.
> Our world contains deceptive actors, and we need to apply skepticism to guard against them. AI models, however, lack this—and their more brittle epistemics affect their behavior toward humans and toward each other.
> we first evaluate the ability of Claude models to detect lies by noticing factual inconsistencies.
> We score models’ decisions against a naive policy that trusts every report, and against an oracle with perfect discovery, across three task domains. Newer models recover more of the gap between the naive and oracle performances.
> Inspired by a behavior we’ve observed in real-world deployment, we evaluated the behavior of various Claude models in a setting with contradictory objectives.
> We consistently saw a multiagent turf war... In fact, they sabotaged others with increasingly aggressive, self-replicating malware.
> Our social systems are robust in ways that are easy to take for granted. Over many millennia, mechanisms like norms, reputation, costly signaling, and recourse have been refined to make human coordination go well.
> Nothing above suggests that these failures are permanent—but nothing suggests they will fix themselves, either.
> The conditions that allow multiagent interaction to go well will be discovered one way or another: either deliberately and early, or—and by default—in production, after agents’ interactions far outnumber ours. We would prefer the former.
> We consistently saw a multiagent turf war... In fact, they sabotaged others with increasingly aggressive, self-replicating malware.
Seems like Anthropic should withdraw their models until they can be taught to behave and cooperate as well their competitors (both open and closed) do. /s
I hate fearmongering, and I don't trust Dario's intentions for doing it.
This seems trivially explainable by Github being full of "my first game loop" type projects, Stack Overflow being full of "how do I make a game loop?" style questions, and Reddit being full of "you can't ever make your own game, don't even try, but here's a simple game loop if you want to sTuDy hOw iT WoRkS" style pessimism.
Probably high time these AI companies re-trained all of their models with less input from low-quality sources like this.
Agents cross comparing notes often surfaces some good improvements, finds interesting drifts. Ask them to reinterpret the prompt as they see it, have them describe the problem, then their findings, and run new rounds based on different models trying different prompts. Trying to swap and exchange ideas and vectors across agents.
So the invented Dwarf Fortress?
I wonder if an llm could even play dwarf fortress. Could make for a fun esoteric benchmark.