LLM Agent Paper List
github.com
From the repo - "We start by the general conceptual framework for LLM-based agents: comprising three main components: brain, perception, and action, and the framework can be tailored to suit different applications."
github.com
From the repo - "We start by the general conceptual framework for LLM-based agents: comprising three main components: brain, perception, and action, and the framework can be tailored to suit different applications."
I've taken "Show HN" out of the title now.
I’m becoming convinced Agents are mostly novelty. They aren’t Turing complete. They’re fun to implement, but any algorithm they run to solve a task can’t be inspected and can go awry with complex logic. You can accomplish the same thing by asking the LLM to write code in whatever language, and it’s more reliable. There is more reluctance to do that though, because the idea of a BE environment that runs user generated scripts is scary.
I too have been playing with agents. But I’ve explicitly biased away from having them do things that can achieved directly with code.
Stringing together tool selection into steps with reasoning imo is the crucial brain layer. Code is one of the tools. I’ve also gone with a kind of preemptive memory, where the brain gets injected with relevant memory automatically, instead of it requesting for it. Same for retrieval from the internet.
But wanted to understand how you’re thinking about.
I think most of the actual value created short/medium-term with agents will be in entertainment where things going awry is OK and no action they take can really be catastrophic.
Some people are already spending a ton of time on character.ai (https://beta.character.ai). There are sims-like games like ai-town (https://www.convex.dev/ai-town). My friend and I made an open-source murder mystery novella with agents (https://gron.games).
We found that a lot of the actions they take are pretty low stakes, so the correctness issues are much less important (in ai-town, they just talk with each other, in our game, they just beat you or run away if you interrogate them poorly).
I think this is why agents appear suited to entertainment, but this is almost a different topic that I could rant about for hours. There’s a whimsical element of randomness in what the character may do that the programmer does not need to explicitly design. However, I think that whimsy has limited reach, so while the developer may be more amused by their own creation it has limited appeal.
[not a comment on OP post]
People have been thinking of a way to make LLMs perform actions in the real world, by giving these LLMs “tools”.
An LLM with tools that can perform tasks would be an AI Agent
What makes it different from other programs is that Agents can do probabilistic inference on the closure natural language, i.e. they can predict(P(Y|X) as well as generate(P(X, y)). Both of these are very hard to do with any one previous model and hence with any service that could combine software API + Classic ML Model(pre Generative+Discriminative models like llm).
Today's LLMs are kind of like Dory from Finding Nemo - you have to recreate the context every time you do a slightly different task, or when the context window for the LLM is no longer sufficient to remember previous turns in a conversation.
An agent can sit on the other side of a piece of collaborative software that was designed for human collaboration. This is why chatbots are the current "killer app" for AI, we already understand how to collaborate with other people via chat.
Now imagine what we can do with more sophisticated pieces of software like Figma or Excel. Disclosure: I work on the Python in Excel feature and we just announced Copilot for Excel a couple of days ago. [2] Other modalities that excite me are tools like Vision Pro which will be an interesting test bed for multi-modal generative agents.
source is available if you tap the GH icon in the footer.
The application uses an LLM to breakdown the problem into pieces and asks the LLM to choose a tool that is suited to handle each piece. The application then invokes the tool to solve that piece of the problem. This is iteratively done until the original problem is solved or determined to be unsolvable.
There are many agent variations, depending on whether you breakdown the problem once at the beginning, or understand what to do next based on the outcome of the previous step, or constantly evaluate and reprioritize which piece of the problem to solve next.
Because you rely on an LLM to understand the problem, if you chose an LLM that exhibits good world understanding, and designed the LLM Agent well enough, then you can use it to solve problems that traditionally required deterministic code.
You can augment it with retrieval systems(vector store, graph db, sql db + lexical/semantic/graph search) and give it access to tools like api, search engine and code interpreter. Then it becomes an agent.
Agent= Reasoning + Memory + Planning + Tools
I wrote this article recently that outlines what a system needs to run agents: https://github.com/lukebuehler/agent-os/blob/main/docs/artic...
You can read a good intro example here: https://github.com/brexhq/prompt-engineering#react
1. Yeah, I know HN's feeling about discord and posted this anyway.
There are 265 papers currently.
If you're curious, there's a reason it feels overwhelming. Here's a plot of the age of the papers relative to now. Its a bit of a wall/tsunami.
https://i.imgur.com/Om7udZR.png
Edit: Because of the shock like character I also postulate several hypotheses:
- It will rapidly expand like a fireball, consuming all of the "air" of possible nearby topics.
- Possible secondary: Ignition of other topics as possible fuel to extend the publishing "burn."
- Possible secondary: Vacuum implosion when all nearby topic fuel is expended.
- Cause significant destruction (existing wealth, businesses, market segments, ect...)- Cause many humans to effectively "reflect" due to its impenetrability, effectively rendering those humans permanently downhill as consumers.
- Possible secondary: Further shocks as the shock wave bends or reflects off of "currently impenetrable" obstructions (topics that do not immediately turn into fuel).
- Cause significant secondary (mostly unrelated) growth in topics impeded by the current suffocating fuel environment (like an old growth forest that has become choked with debris and oppresses all growth)