[not a comment on OP post]
[not a comment on OP post]
Today's LLMs are kind of like Dory from Finding Nemo - you have to recreate the context every time you do a slightly different task, or when the context window for the LLM is no longer sufficient to remember previous turns in a conversation.
An agent can sit on the other side of a piece of collaborative software that was designed for human collaboration. This is why chatbots are the current "killer app" for AI, we already understand how to collaborate with other people via chat.
Now imagine what we can do with more sophisticated pieces of software like Figma or Excel. Disclosure: I work on the Python in Excel feature and we just announced Copilot for Excel a couple of days ago. [2] Other modalities that excite me are tools like Vision Pro which will be an interesting test bed for multi-modal generative agents.
You can augment it with retrieval systems(vector store, graph db, sql db + lexical/semantic/graph search) and give it access to tools like api, search engine and code interpreter. Then it becomes an agent.
What makes it different from other programs is that Agents can do probabilistic inference on the closure natural language, i.e. they can predict(P(Y|X) as well as generate(P(X, y)). Both of these are very hard to do with any one previous model and hence with any service that could combine software API + Classic ML Model(pre Generative+Discriminative models like llm).
I wrote this article recently that outlines what a system needs to run agents: https://github.com/lukebuehler/agent-os/blob/main/docs/artic...
The application uses an LLM to breakdown the problem into pieces and asks the LLM to choose a tool that is suited to handle each piece. The application then invokes the tool to solve that piece of the problem. This is iteratively done until the original problem is solved or determined to be unsolvable.
There are many agent variations, depending on whether you breakdown the problem once at the beginning, or understand what to do next based on the outcome of the previous step, or constantly evaluate and reprioritize which piece of the problem to solve next.
Because you rely on an LLM to understand the problem, if you chose an LLM that exhibits good world understanding, and designed the LLM Agent well enough, then you can use it to solve problems that traditionally required deterministic code.
You can read a good intro example here: https://github.com/brexhq/prompt-engineering#react
Agent= Reasoning + Memory + Planning + Tools
People have been thinking of a way to make LLMs perform actions in the real world, by giving these LLMs “tools”.
An LLM with tools that can perform tasks would be an AI Agent
source is available if you tap the GH icon in the footer.