MetaGPT: Meta Programming for Multi-Agent Collaborative Framework
arxiv.org
arxiv.org
I mean that intuitively I couldn't imagine replacing 1 experienced professional with 1, 2, 10, 100 or even 1000 intelligent high school graduates. Intelligence doesn't seem strictly additive across multiple individuals for all cases. But then I consider that one of the most powerful life-lines in the TV game show "Who wants to be a millionaire" was the "Ask the audience". I am reminded of the cliche of the wisdom of crowds, even when the crowd is made up of non-experts.
This suggests to me that there is a kind of problem where multiple lower power agents can solve the issue to a higher quality. But there are also kinds of problem where a single higher-power intelligence will be necessary. I haven't developed an intuition when each approach is valid.
Architecture as in you may have multiple GPT instances being prompted to look at a problem from different angles: "Analyse the problem as a pessimist", "As an optimist", "as a mathematician", "an engineer", "a philosopher", etc.
Caching, as in determining what you can store from these outputs, e.g. "give a numeric 1-10 score of what you think of this product as a LGBT-friendly conservative from the Midwest", etc.
Sure, but just as a thought experiment imagine you have a really bad disease. I give you two options: 100 high-school students can diagnose and prescribe a treatment or you can choose 1 professional with 10 years of experience in related diseases.
You might initially prefer the professional with 10 years of experience. Would your opinion change if I told you that I selected the high-schoolers to be diverse so that one is pessimist, one is an optimist, one got good grades in engineering, one loves philosophy?
Of course, my intuition might be wrong. For example, perhaps it legitimately would be better to have 100 high-school students where one is a high-school level ability with ear-nose-throat, one is a high-school level ability in oncology, one is a high-school level ability in cardiology, etc. Except some control AI would have to synthesize their answers into a coherent response ... and that controller would be high school level.
I'm not really sold either way if I am honest. I don't think we have the answers to these questions. It just shows that my intuition about intelligence is open to challenges.
What I mean to say is, if I am considering building a product based on LLMs then I may have to make a basic decision: can I use multiple cheap LLMs in a multi-agent setup or must I use a single expensive powerful LLM. Right now I don't have any intuition on what kinds of problems are solved most efficiently by either approach. Just looking at a problem description I can't intuit which approach is appropriate.
If you try to start with the multiple cheap LLMs approach, it is almost certainly going to be more difficult to get the prompting right, and if you don't already know what you are trying to accomplish is feasible, you're adding a lot of work that might add up to nothing (even if it would actually work if you used the more powerful model).
Meanwhile, the "100 high school students" on a mobile phone are the only medical consultation someone in Sub-Saharan Africa or Southern Asia is going to have access to at all.
"Why Not Both" would surely apply here. Maybe we can't give everyone access to a kind, caring, patient human doctor; but we sure can make coming in second place a lot less painful.
[0]https://www.who.int/news/item/13-12-2017-world-bank-and-who-...
It will happily hallucinate facts and claim they are true. Well, ok, maybe it is ;)
The issue they seem to tackle is trying to minimize hallucinations by injecting some human expertise into the pipeline. They do this by more strictly defining the roles and tasks a step in that pipeline needs to accomplish.
As if humans don't, bullshitting:
https://www.sciencedirect.com/science/article/abs/pii/S00221...
https://www.poynter.org/fact-checking/2018/this-study-is-all...
I see two camps, people quietly working, and people working on grandiose ideas from the initial rush that make good headlines but not good products.
Some informal warning signs I use subconciously:
- "paper" on Arxiv, and its about prompt engineering
- agent_S_
- state machine where the LLM is eating its own output repeatedly
- LLM makes decisions based on its own output.
There's 100x more alpha in the obvious stuff because even that's not well-implemented or shared widely. Ex. people are still stuck on hallucinations 90% of the time: an obvious way to handle that is doc retrieval.
Last n.b.: in 2021 I was frustrated with ~100% of models being behind locked doors. Then, I saw how GPT3.0 was working and evolving. My mantra became "products not papers." Maybe that applies downstream of the LLM now.
I agree, in the sense that my intuition is fed by having as many available examples as possible. I am not even sure one could quantify the difference between intelligence levels in a way that I would be satisfied with in a paper.
For example, I find it intuitive that a mixture of 10 narrowly fine-tuned GPT-3s are better at a task than 1 broadly fine-tuned GPT-3. But I don't have a real intuition about how many GPT-3s you would have to mix to match the quality of GPT-4, or if there even is any number of GPT-3s you can mix to achieve the result of GPT-4. I think we just need to start building systems and see what happens.
We tried it recently and found it very challenging to get it to accomplish simple tasks.
re: AutoGPT
Got really excited at first. Thought maybe I had missed that it was viable in a year of playing with LLMs. Tried a few demos, didn't work. Looked into it more and confirmed a core loop involved LLM eating its own output to make decision.
I still kept investigating it on my todo list, in case my earlier experiments with that approach were wrong.
I took it off the todo list later.
I saw near-universal feedback like yours, that general technique never worked IMHO, and IMHO sycophancy explains why. Paper here[1], TL;DR the model is very likely to agree, so critical feedback loops over multiple steps tend to settle into a loops of steps.
IMHO this doesn't mean sycophancy breaks _all_ workflows, ex. a flow for writing a story involving outlining, writing, criticizing, then rewriting is a genuine real quality boost.
However, if a human does write => criticize 10 times, it keeps getting better each iteration. If you have an LLM do it 10 times, IMHO it's actively harmful after round 3.
"outline" => ["write page 1", "write page 2", "write page 3"] => ["feedback on page 1"..."feedback on page3" => ["use feedback and original draft to rewrite page1"..."page3"] => "combine pages 1 2 and 3 into cohesive story"
[1] https://www.anthropic.com/index/discovering-language-model-b...
It's not that it can't improve the writing further per se, but that it takes very detailed prompting to get it to give a detailed enough critique to do so consistently enough across even a page (e.g. ot might come uo qith a great lone but proceed to edit out the best paragraph elsewhere) to the point that I tend to agree with you in as much as it at least will take a much more convoluted chain of prompts to maybe get there at the moment, and you'll be fighting GPT4s tendency to actively cheer on really juvenile prose the whole way.
In a way I think the biggest hindrance to get it to write better at the moment is that it has awful "taste", and having to explicitly give it a long list of rules to check against is a poor substitute.
As an aide, though, I think you could get reasonable but not great results at "bridge these two paragraphs and maintain the style" type tasks, or expanding descriptions into a paragraph or two, though more so for non-fiction writing.
For fleshing out the basics of a technical spec and pointing out what I've missed I've had decent luck, on the other hand. It's not come up with any Earth shattering revelations, but for a dry spec that's not the point.
Telephone (game) https://en.wikipedia.org/wiki/Telephone_(game)#Game :
>> The game has no winner: the entertainment comes from comparing the original and final messages. Intermediate messages may also be compared; some messages will become unrecognizable after only a few steps.
Transmission chain method: https://en.wikipedia.org/wiki/Transmission_chain_method :
>> The transmission chain method is a method used in cultural evolution research to uncover biases in cultural transmission.[1] This method was first developed by Frederic Bartlett in 1932.[2][1]
Feedback > Electrical Engineering; Positive Feedback, Negative Feedback,: https://en.wikipedia.org/wiki/Feedback#Electronic_engineerin...
Control Theory > Stability: https://en.wikipedia.org/wiki/Control_theory#Stability
Multi-agent System > Concept: https://en.wikipedia.org/wiki/Multi-agent_system#Concept
Convergence (disambiguation) https://en.wikipedia.org/wiki/Convergence
Convergence (logic) https://en.wikipedia.org/wiki/Convergence_(logic) :
> In mathematics, computer science and logic, convergence is the idea that different sequences of transformations come to a conclusion in a finite amount of time (the transformations are terminating), and that the conclusion reached is independent of the path taken to get to it (they are confluent).
> More formally, a preordered set of term rewriting transformations are said to be convergent if they are confluent and terminating.[1]
And then
Consensus (disambiguation) https://en.wikipedia.org/wiki/Consensus_(disambiguation)
Revision is a step in a Writing Process.
Revision (writing) https://en.wikipedia.org/wiki/Revision_(writing)
Writing Process: https://en.wikipedia.org/wiki/Writing_process
Collaborative writing > Types, Tools: https://en.wikipedia.org/wiki/Collaborative_writing#Tools
And now it is time to underline the premise(s) and the conclusion(s), time to apply Critical Thinking; Logic and Rationality.
Critical Thinking > Logic and rationality: https://en.wikipedia.org/wiki/Critical_thinking#Logic_and_ra...
To simulate a full-scale multi-agent game, there would need to be a "Fourth Estate" (and maybe a Fifth Estate); a peanut gallery of dissenters with signs and no jobs.
Then, you could model consensus with LLMs and have something better than the mediocrity that sometimes results from committees.
More samples from the same LLM vs Sample different LLMs
Maybe feed one or more agents relevant encyclopedia articles as context first; which read the most encyclopedia articles first?
It's worth noting that GPT-4 internally uses a Mixture of Experts (MoE) model with 8 'experts' internally, so it's more similar to a multi-agent setup than you might think initially.
Has this actually been confirmed, either officially by OpenAI or otherwise? As far as I know, George Hotz claimed this once, and since then everyone just assumed it was the truth without actually waiting for any sort of verification.
My naive understanding of layers in a model is that each layer loosely acts as an expert in one step of the entire process.
For example, in an object recognition model, one layer takes on the task of separating objects from the background, another excels at knowing the colour of different things, another might learn the difference between a blue sky vs. the colour sky blue.
So essentially, we’re trying to mimic the same working model at a higher level of abstraction. Similar to how our body is made of atoms. Many atoms make a molecule. Many molecules make organic tissue, and amino acids that perform more complex operations. Skip all the way up and you have a human being. Human beings in numbers can put a man on the moon, create anti-matter and nuclear explosions.
In theory - this can work just as well. One agent to provide a solution, another to critique it, another to verify it, another to mimic the end-user. The big obvious missing piece is memory and the ability to learn while doing (at least in existing LLMs).
It’s like having a software development team that’s frozen in time and knowledge. These LLM agents will always require some micromanagement and hand holding, and will waste a lot of resources with failed attempts - that they are going to repeat every time you ask them to perform this task.
The intuition I am asking for is more like: "Are 10, 100, or 1000 GPT-3s capable of producing the equivalent intelligence of 1 GPT-4". That kind of reasoning expanded to GPT-N.
And further, if no number of GPT-3s can reach the level of intelligence of a GPT-4, then what set of problems can be solved by 100s of GPT-3s in a mixed model and what problems require the single GPT-4.
That's a very common misunderstanding of what MoEs are. What you're describing is an Ensemble.
A MoE is where a the input is "routed" to one of X "experts" in the higher layers of the network. Something like this: {input} -> [lower layers] -> <decide which upper layers to send input to> -> [expert] -> output
And these experts aren't experts in the space of human defined subjects. Their expertise are in some embedding space, which might or might not line up with our intuition of a subject.
Rationale:
- Speed - Ability to recall - Depth and breadth of “exposure”
I can’t speak to others high school experience, but in my experience, high school does not prepare young adults to be useful in the same settings that GPT-4 could be useful.
Example prompt: - Pose this to a high school graduate and GPT-4, somehow benchmark quality.
1. “You are an experienced business strategist. Your client has the following problem <insert problem synopsis here>. Generate an outline of brainstorming topics your client should consider to accomplish <insert goal here>.
- First, GPT-4 is going to respond in ~30sec or less - Second, a high school grad has 0 context for these types of prompts.
Useful for what?
LLMs are largely useless for the vast majority of labor that people with HS degrees tend to do (which mostly involves not being behind a keyboard all day).
This analogy doesn't make sense, because the professional is presumably also a high school graduate. This case is more like leveraging a team of specialists with expertise in different domains.
I see it as a matter of attention. GPT-4 is limited in the number of tokens you can feed it and receive back from it. I see this as how much attention the LLM can give you. I see meta-gpt and other models like this as increasing the attention of the LLM by allowing it combine it's short attention span with multiple assessments of the request, and give you a more complete picture of what you want because it can keep giving it consistent attention to the problem at hand, instead of simply trying to solve it on the first try.
Someone please correct me if I'm way off, but this simple mental model helps me abstract this.
You can use multiple agents, or split a lot of information across multiple requests to one agent. The result is the same. Some problems require a full understanding of the whole picture.
Humans love to think of multi agent systems as being like a team of people. It's much more like a writer imagining different characters and how they would respond. When George RR Martin imagines all 500 characters in Game of Thrones, there is a lot of diversity of perspective and thought there. But all of that is coming from one intelligence and doesn't represent a collaboration in any traditional sense.
No, its not.
Its more like 1,000 individuals sharing one genetic template.
Its just that the experiential context that along with the template makes an individual is very small with LLM instances.
No it's not. It does no good for a Language Model to configure a global persona. It needs to be able to predict text from wildly varying backgrounds and contexts. It's not pretending anymore than anything else it does is pretending.
That's why experiments like the below actually work
Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus? (https://arxiv.org/abs/2301.07543)
Out of One, Many: Using Language Models to Simulate Human Samples (https://arxiv.org/abs/2209.06899)
A perfect LLM would predict Einstein as well as it would predict the dumbass down the street.
Now RLHF does incentivize a more global persona by default but stepping away from that is trivial
The phenomenon I am describing is known as "Emergence."
The creation process used 11,940 tokens on input and 2,993 tokens on output, which cost $0.35 and $0.18, respectively.
The game it generated consisted of four python classes in four separate files: Main, Game, Snake, and Food.
The game executed without error on the first try, but the snake wasn't able to 'eat' the food. Here's the relevant code for 'eating' food:
# Check if the snake ate the food
if self.snake.body[0] == self.food.position:
self.score += 1
self.snake.grow()
self.food.generate()
The issue was that the snake's body was represented as a list of lists, whereas the food position was stored in a tuple. After changing the food position to a list, the game worked correctly.https://chat.openai.com/share/b4b399ef-1def-4f68-b2f1-8c56ca...
Seems to work correctly, didn't have to change anything in the code.
Like all the other agenty stuff I've seen it's not clear what the fluff adds over just prompting the base model.
"Write 10 popular mini games"
--> What are those games? 1. Agent --> How does each of them work? 2. Agent --> Write each of those. Agent 3-13
"Implement a Gomoku game using Python, incorporating an AI opponent with varying difficulty levels."
One shot prompt works fine: https://chat.openai.com/share/4435152f-ba13-45a5-a116-eb4f57...
AFAICT using ChatGPT directly gets you better results with lower cost and less complexity. The "MetaGPT" framework adds bloat and delivers nothing.
https://chat.openai.com/share/561419c8-4143-4172-b5e2-a411ee...
I'm not convinced the meta/agenty/company stuff is doing anything to help the LLM generate working code and it looks like they never bothered to check the null hypohesis before hitting publish.
Personally, I think the use cases that go beyond chat are going to be the most valuable. Specifically, the ability to produce structured information.
It's easy to imagine how systems like these might succeed by following workflows analogous to those humans use, on greenfield toy problems.
The trouble is, if you paratrooper something like this into an exiting repository, or you want to build something of real significance, then you've got more context than can fit into an LLMs window. You don't just have the hallucination / stay-on-task problems, you have the problem of the LLM not having everything it needs to know to complete its task.
In my experience, the latter problem of ensuring the LLM knows everything it needs to know is the bigger problem. The other stuff (using standard operating procedures and workflows, as in this paper) is actually the relatively easy stuff. I'm not saying it's easy-easy or that it's obvious from the outside if you haven't played around with it; I'm not disparaging the paper.
So, what I'm most curious about, and I wonder if anyone here has seen good solutions, is this:
How do these multi-agent systems maintain/retrieve the appropriate domain context required to make good design decisions, good coding decisions, and so on? "Use embedding and shove stuff into a vector database" is hand-wavy and doesn't get you all the way there. I'm looking for concrete solutions; academic papers that show a lot of promise; and the like.
(Maybe they lay it out in this paper and I just haven't gotten to it yet. But if they've cracked that nut I suspect they would lead with it.)
Yes. Not everyone will admit as most big companies have policies against it. Small companies are, because the cost saving is so high.
Feel free to answer only subset of the questions that you feel comfortable with them.
our entire focus is on multi-agent workflows based and trained on real peoples published work. I think the output is amazing, but I'm biased of course. Its ability to accomplish soft tasks like advice, help etc is far superior to vanilla GPT4
> ## Code: {filename} Write code with triple quoto, based on the following list and context.
> 1. Do your best to implement THIS ONLY ONE FILE. ONLY USE EXISTING API. IF NO API, IMPLEMENT IT.
1. https://github.com/geekan/MetaGPT/blob/main/metagpt/actions/...
The future will likely be simulated companies and the business logic will be how you compose this agent interaction.
I haven't run it yet but I feel this specific paper is just good prompt engineering (with a little bit real engineering in it).
Also it reads like a pitch honestly. Would bet they are looking for funding rn.
This was a take-home challenge I got back in 2015 for an internship. I spent all day on it and had a blast. Curious to compare my nooby college code to state of the art LLM code!
It's pretty interesting how well it works, but there is a lot more to improve as you can refine the agents for more than their generalized domain, into their perspective and real publish knowledge for increased improvement.
That being said, totally understood. Logging in is very little cost though, we just need to make a better value prop upfront.
Anyways nice to see some discussion here, didn't read a lot elsewhere about the project.