Build your own agents which are controlled by LLMs
github.com
github.com
That said, the bit about "The trick to avoid hallucination" in the attached blog post is a very neat idea. Obvious in hindsight, as it often is:
> So the trick is that we send a stop pattern which in this case is when we see Observation: in it’s output, because then it has created a Thought, an Action and used a tool and hallucinates the Observation: itself :D
> This stop parameter is a normal parameter of the OpenAI API by the way, so nothing special to implement there.
In general, it's magical when you put these things in a feedback loop. The way they can automatically respond to errors and adjust their actions is really cool - take a look at the last gif here[1].
It worked quite well. To well almost: I started a meta-conversation where I asked another GPT4 instance to come up with conversations SoulverGPT could have with a user where the addition of solving is beneficial. This worked, and eventually even found a bug in Soulver - essentially fuzzing the language.
[1] https://comby.dev/blog/2022/04/11/comby-decomposer-compiler-...
Really neat to see the parallels to OODA loops, which are a frequent feature (implicit or explicit) of human decision-making processes.
I had never heard of those, or PDCA before, but I'm working on a startup in this space (AI Agents), and found PDCA by just using bing's suggestions i.e. OODA loops vs ...
Anyhow, I think PDCA looks more promising actually, a little more logical/analytical IMHO.... but don't take my word for it, here's what gpt4 had to say :
Which would be a better framework for AI agents working in steps and loops to complete tasks? OODA or PDCA?
Both the OODA loop and PDCA (Plan, Do, Check, Act) cycle can be applied as frameworks for AI agents to go through retrospection phases or steps while performing tasks. The choice between the two depends on the specific context and the desired outcome of the retrospection process. Let's examine both frameworks in this context:
OODA loop:
Observe: The AI agent gathers information about its current task, the environment, and any potential obstacles.
Orient: The AI agent processes and analyzes the collected information, updates its understanding, and identifies possible actions.
Decide: The AI agent chooses the best course of action based on the available information.
Act: The AI agent executes the chosen action, observes the results, and adjusts its approach as needed.
In this context, the OODA loop is well-suited for situations where the AI agent needs to make rapid decisions and adapt quickly to changing circumstances, such as real-time or highly dynamic environments.
PDCA cycle:
Plan: The AI agent develops a plan or strategy for accomplishing its task, based on its knowledge and understanding of the problem.
Do: The AI agent executes the plan, carrying out the necessary actions to complete the task.
Check: The AI agent evaluates the results of its actions, comparing them to the desired outcome or goal.
Act: The AI agent makes adjustments or improvements to its plan based on the results of the evaluation, and then iterates through the PDCA cycle again.
In this context, the PDCA cycle is more focused on continuous improvement and learning, making it a better fit for situations where the AI agent has a clearly defined goal or objective and can iteratively refine its approach over time.
Ultimately, the choice between the OODA loop and PDCA cycle depends on the specific requirements and goals of your AI agent. You may even find that a hybrid approach, combining elements from both frameworks, is the most effective solution for your AI agent's retrospection process.Mostly people just see it as a circle in a PowerPoint slide and it has close to zero utility in that form.
To be fair to langchain, though: I like it a lot and it offers a lot more than my repo, but yeah definitely not the best place to understand how it all plays together.
That said, while I am very, very impressed with GPT4 for lots of uses right now, so far it's just not clear to me that feeding it back into itself is fruitful at this point.
When I use GPT4 for coding, I give it MUCH higher level instructions than I use on a search engine, much closer to the actual problem I'm solving. But I'm still breaking the problem down into smaller problems; I need to read the output, fix errors or instruct it to fix errors, and then have it build more features on. It's similar with creative processes, brainstorming, and other writing.
These agents largely strike me as an attempt to replace this whole fact-checking/editing type routine with the LLM itself; but seeing as it's the thing the LLM is not yet good at, I'm not sure how much progress can be made there, vs just waiting for GPT5 and hoping it's another big leap in capabilities.
Just not sure feeding it back to itself does much at the moment.
If it 0-shots something and executes it and halts, not an agent, if it observes output and loops once, that's a minimal agent. I'm not super skeptical of a couple loops; I just don't know that GPT4 has what it takes to drive a larger number of loops.
If I had GPT-4 API access, it would have found that on the first try. Sigh.
--
I've noticed it does skip the thought process sometimes.
Question: Who is the president? What year whas he born. Name one other famous person born in that year. Thought:
Final answer is The current president of the United States is Joe Biden, born in 1942. Some other famous people born in the same year include Harrison Ford, Muhammad Ali, and Aretha Franklin.
So even though there was a thought, I didn't print it
Some kind of ensemble agent which is more robust, might play with that idea.
Though the project makes a decent GUI to use ChatGPT if anyone is interested: https://github.com/blipk/gptroles
You can run and edit code snippets in the chat interface
Maybe another Winter was staved off by the heat of a million GPU suns (energy consumption notwithstanding).
In the future, we'll spray our GPUs with little spritzes from a water bottle, because something-something harvesting of photosynthetic electron leakage.