Emerging architectures for LLM applications
a16z.com
a16z.com
That said, while much of this might not have any real traction long-term, looking at what researchers use seems to miss the mark a bit. It’s like saying network technology researchers aren’t using Vercel.
I really appreciate that they called out and separated some hype vs. practice, specifically with regards to Agents. This is something I keep hoping works better than it does, and in practice every attempt I've taken in this direction leads to disappointment.
I feel like lots of paper are getting published and reviewed which is good as bad ideas don’t get to propagate for ages
We've spun off a company to realize the vision of bringing an open-source, standardized framework to the PHP ecosystem, where we've been building apps for communities for over a decade. It's "AI for the rest of us", but at the same time promoting positive collaboration between communities and AI. It also involves micropayments for tasks done by either an AI or human agent.
If you're a VC or an expert in the space, I'd love to get feedback on this: https://engageusers.ai/ecosystem.pdf
And if you want to get involved in any capacity, whether as an investor or developer, please email me greg at the domain engageusers.ai -- this time around we are planning to take on venture capital funding for this project, and syndicate a round later this summer.
pip install milvus
Other than that, it's great to see the shout-out for Vespa.Vector database space is the Wild West, keep at it
I think this is a very well articulated breakdown of the "LLM Core Code Shell" (https://www.latent.space/p/function-agents#%C2%A7llm-core-co...) view of the world. but it is underselling the potential to leave the agents stuff to a three paragraph "what about agents?" piece at the end. the emerging architecture of "Code Core, LLM Shell" decentralizing and specializing the role of the LLM will hopefully get more airtime in the december a16z landscape chart!
we actually just purposefully left that part a bit scarce because we have something else coming up on the topic! I'm sure we will be chatting through it soon :)
I do feel the article presents old concepts as "emerging".
if you are curious about building something quickly, you can jump into one of the tutorials https://haystack.deepset.ai/tutorials
Over a weekend I've used deepset/haystack to build a Q/A engine over open source communities slack and discord threads that can potentially have an answer - it was a joy and a breeze to implement. If you have question about Metaflow, K8s, Golang, Deepset, Deep Java Library and some other tech - try asking your quick question on https://www.kwq.ai :-)
I see all the same bits called out:
- Data collection
- Machine /resource management
- Serving
- Monitoring
- Analysis
- Process mgt
- Data verification
There's some new concepts that aren't quite captured in the original paper like the "playground" though.
I've kind of been expecting a follow-up that shows an update to that original paper.
[1] https://proceedings.neurips.cc/paper_files/paper/2015/file/8...
I see sub second performance on >1m vectors with PgVector. Vector databases have a place but this statement seems disingenuous at best. Bringing on a vector database adds additional complexity and a giant chunk of use cases simply don't need it. Not to mention the additional latency you'd be adding.
I don't even think this is a correct definition for "in-context learning". In-context learning is a type of few-shot learning in which examples of input/output pairs are provided as part of the prompt. The idea is that the model is able to "learn" the pattern of the task from the examples. Quoting from the GPT-3 paper:
>what we call “in-context learning”, using the text input of a pretrained language model as a form of task specification: the model is conditioned on a natural language instruction and/or a few demonstrations of the task and is then expected to complete further instances of the task simply by predicting what comes next.
I really don't think it's standard to refer to the process of embedding-based retrieval as "in-context learning".
Only thing special is that the input for each demonstration is obtained through embedding-based retrieval.
Shameless plug: For fellow Ruby-ists we're building an orchestration layer for building LLM applications, inspired by the original, Langchain.rb: https://github.com/andreibondarev/langchainrb
Many are frustrated about not being able to better direct the agents
It’s like the agents have certain pre-learned things they can do, but they aren’t really learning how to apply those things to the environments their human operators want them to develop in
Or at least it is not easy/straightforward how to teach the model new tricks
The Agent & Tools metaphor makes a ton of sense for its simplicity, but has yet to scale beyond very simple agents.
I’m long on it as a programming model, but a lot of work is needed.
Enterprise, not so quite, there are a lot of other stuff to consider like points missing, ethical application, filtering, security, points that are very important for enterprise customers.
Also, in-context learning is just one way to apply LLMs, there are way more applications like few short learning, fine tuning, depending on cost and application involved as I've highlighted here
Other problem with agent is : most independent agents are capable of doing very thin slice of use case, but for complex knowledge work tasks, more often than not, one agent is not enough. You need a team of agents. We introduced a concept of Agent Clusters - which operate in master slave architecture and coordinating among themselves to complete nuanced tasks and coordinating via shared memory and shared task list.
Another big bottleneck I think is lack of a notional concept of Knowledge for Agents. We have LTM and STM, but knowledge is specialized understanding of particular class of objectives ( ecommerce customer support, Account based marketing, medical diagnostics for particular condition etc ) plugged into the agent. Currently agents leverage on the knowledge available in the LLMs. LLMs are great for intelligence, but not necessarily knowledge required for an objective. So we added concept of knowledge - which is a embedding plugged into agent apart from LTM / STM
There lot of other challenges that need to solved like agent performance monitoring, agent specific models, agent to agent communication etc to truly solve for agents deployed in production. Not sure about point mentioned in the article that they might even take over the entire stack because autonomous agentic behaviour is good for certain use cases and not for all kinds of apps.
[1] https://github.com/terminusdb-labs/terminusdb-semantic-index...
The space of possible different "graphs" of LLM agents connected to each other is even larger.
Each graph represents a multiagent system
Here's a generic syntax for notating graphs that don't have loops (essentially trees):
AgentName: Descriptive Name Goals: - Goal1 - Goal2 ...
Techniques: - Instruction1 - Instruction2 ...
Inputs: - From AgentName: Description of input - From OtherAgentName: Description of input ...
Outputs: - To AgentName: Description of output - To OtherAgentName: Description of output ...
-> SubAgentName1: Descriptive Name Goals: - Goal1 - Goal2 ...
Techniques: - Instruction1 - Instruction2 ...
Inputs: - From AgentName: Description of input - From OtherAgentName: Description of input ...
Outputs: - To AgentName: Description of output - To OtherAgentName: Description of output ...
-> SubSubAgentName1: Descriptive Name
Goals:
- Goal1
- Goal2
Techniques:
- Instruction1
- Instruction2
Inputs:
- From SubAgentName1: Description of input
Outputs:
- To SubAgentName1: Description of output
...
-> SubSubAgentName2: Descriptive Name
Goals:
- Goal1
- Goal2
Techniques:
- Instruction1
- Instruction2
Inputs:
- From SubAgentName1: Description of input
Outputs:
- To SubAgentName1: Description of output
...
-> SubAgentName2: Descriptive Name
Goals:
- Goal1
- Goal2
...Techniques: - Instruction1 - Instruction2 ...
Inputs: - From AgentName: Description of input ...
Outputs: - To AgentName: Description of output ...
---
Here's an example researcher agent and it's interior using the syntax. The English translation was lost in my notes but you can put this to a translator:
Tutkimus: Tavoitteet: - Tuottaa ja analysoida tiedusteluja - Rakentaa uutta tutkimusta - Tutkia tuntemattomia aiheita - Päätellä tiedusteluista Tekniset ohjeet: - Ohje / sääntö päättelylle 1 - Ohje / sääntö päättelylle 2 Syötteet: - Agentilta Meta-tietoisuus: Ehdotukset - Agentilta Alatutkimus: Tulokset Tulosteet: - Agentille Alatutkimus: Käskyt - Agentille Muisti: Tutkimus
-> Alatutkimus: Tavoitteet: - Suorittaa erityisiä tutkimustehtäviä Tekniset ohjeet: - Ohje / sääntö käskyjen noudattamiselle 1 - Ohje / sääntö käskyjen noudattamiselle 2 Syötteet: - Agentilta Tutkimus: Käskyt Tulosteet: - Agentille Tutkimus: Tulokset
-> Muisti: Tavoitteet: - Ylläpitää tutkimuksen tallennetta Tekniikat: - Ohje / sääntö muistikantojen käyttämiselle 1 - Ohje / sääntö muistikantojen käyttämiselle 2 Syötteet: - Agentilta Tutkimus: Tutkimus Tulosteet: Ei mitään
---
Signal theory becomes relevant when thinking about I/O, embedded agency, and when the agents aren't / cannot be constantly "reading" each other.
---
For similar projects, the current AutoGPT-style systems are very primitive and haven't adapted to my ideas. If what I call the cognitive architectures of the LLM-multiagent systems were carefully designed, which I predict will become a thing (and subject to ton of future research!), our AI systems could gain very advanced cognitive capabilities, perhaps even approaching humans but in their own, formal manner.
One person suggested me this:
https://princeton-nlp.github.io/SocraticAI/
I haven't read it but seems to have similarities.
FWIW I asked because I'm working on a toolkit for applying Monte Carlo tree search to agent graph generation and am always on the lookout for fundamental insights that could help direct its development.
I’ll say this post is rather shallow to be considered technical or even fit the title
Just here to say that you can quickly build a robust feature with only OpenAI's APIs, redis, a text file for the prompt you parameterize (versioned), and a little bit of glue code (no LangChain). You can add instrumentation for observability around that like you would any other code.
I would wager that most enterprise use cases don't need most of the tools listed in this article, and using them is complete overkill.
We are talking about calling an API here people. Maybe what is behind the API is magical and powerful seeming but it's just an API that takes some context, the tokens to generate and a temperature setting.
Doing everything in that diagram is probably overkill for most uses. But using it as a starting point and trimming what you don't need will help a bit with triping over "oops I forgot to include that".
But caching, for instance, doesn't need to be it's own lib does it? I don't want 'semantic caching' I want to cache the exact same query and I can do that without being LLM specific:
from joblib import Memory
@memory.cache
def call_chat_completion_api_cached(max_tokens, messages,temperature):
...
I mean, I guess then I might want to store that somewhere central like redis and maybe slowly I need a specific cache tool. So I get your point. It's helpful to see the possibilities of approaching these problems.But it also does feel like an land grab of supporting libs and infrastructure.
A browser database and React is all I have needed for my LLM apps.
But there's usually a lot of reasons why these architectures could be a useful reference. In my project, I host my own trained LLM and one of the cost efficiencies comes from being able to cache at every step along the way. Then there is a large private media hosting consideration.
There is room for all sorts of setups and I kind of liked how the article mapped out some of the common paths.
> Fine-tuned OpenAI models. The Fine-tunes API allows customers to create their own fine-tuned version of the OpenAI models based on the training data that they have uploaded to the service via the Files APIs. The trained fine-tuned models are stored in Azure Storage in the same region, encrypted at rest and logically isolated with their Azure subscription and API credentials. Fine-tuned models can be deleted by the user by calling the DELETE API operation.
My follow-up question is: how is this different from OpenAI's finetuning?
i built my own last weekend. https://github.com/smol-ai/logger
dumps things to json files, or to a log store. all you need for prompt engineering and monitoring really! no VC needed, no DataDog of AI yet
.
> quickly build a robust feature with only
Much like how building Twitter is a weekend-sized project.
More like a month, but yes. That's what we did - a month to launch, then another month to harden and add several features. Rolled out to all customers.
Building product features with LLMs is difficult, but not because of the architectural needs. It's an API you pass data to.
> Claude offers fast inference, GPT-3.5-level accuracy, more customization options for large customers, and up to a 100k context window (though we’ve found accuracy degrades with the length of input).
So yes it does include prompt injection, but is a bit broader. Data Leakage is one that several customers have called out, aka accidentally leakage PII to underlying models when asking them questions about your data.
I'm evaluating tools like Private AI, Arthur AI etc. but they're all fairly nascent.
My email is beady.chap-0f@icloud.com
looks like they are playing catchup and trying to stay relevant.
What happened to their web3 vision ?
Confluent calls it a "customer 360" problem [1] and I don't disagree.
We (Estuary) also wrote up a post showing an approach for Slack => ChatGPT => Google Sheets [2], and have more content coming for Salesforce, HubSpot, and some others.
[1] https://www.confluent.io/blog/chatgpt-and-streaming-data-for...
Good: We're (of course) doing a lot of these architectures behind-the-scenes for louie.ai and client projects around that. Vector embeddings are an easy way to do direct recall for data that's bigger-than-context. As long as the user has a simple question that just needs recalling a text snippet that fairly directly overlaps with the question, vector embeddings are magical. Conversational memory for sharing DB queries across teammates, simple discussion of decades of PDF archives and internal wikis... amazing.
Not so good: What happens when the text data to answer your question isn't a directly semantic search match away? "Why does Team X have so many outages?" => What projects is Team X on" + "Outages for those projects" + "Analysis for outage" . AFAICT, this gets into:
A. Failure: Stick with query -> vector DB -> LLM summary and get the wrong answer over the wrong data
B. AutoGPT: Getting into an autoGPT langchain that iteratively queries the vector DB, and iteratively reasons over results, & iteratively plans, until it finds what it wants. But autoGPT seems to be more excitement than production use. Many questions like speed, cost, & quality...
C. Knowledge graphs: Getting into use the LLM to generate a higher-quality knowledge graph of the data that is more receptive to LLM querying. The above question now becomes a simpler multi-hop query over the KG, so both fast and cost-effective... If you've indexed correctly and taught your LLM to generate the right queries.
(Related: If you're into this kind of topic, we're hiring here to build out these systems + help use them on our customers in investigative areas like cyber, misinfo, & emergency response. See new openings up @ ttps://www.graphistry.com/careers !)
So you have token embeddings, but tokens are too small to be useful.
Is "what a sentence means" encoded as a vector once you have passed the embeddings through a transformer or two?
As a blackbox, it is a generalization of word2vec to sequence2vec. For example, simply summing or averaging word vectors in a sentence can give you a fast & cheap sentence embedding.
But natural language sentences have more structure than natural language words. Ex: it matters precisely where "not" goes in a sentence. So a lot of impressive scientific experimentation went into making these models smarter, with many evolutions. Impressively, this so blackboxed now that doesn't super matter.
Implicit to my post here... that's powerful, and easy to use... but not necessarily a great knowledge representation for someone who wants good Q&A over enterprise-scale data. One of our customer scenarios: "What is known vs believed about incident X." We can index each paragraph as multiple sentence embeddings, so if any phrase matches a query, the full paragraphs can get thrown into GPT as part of our answer. Easy. However, if information in the paragraph may lead to wanting to get information from elsewhere in the system (mention of another team, project, incident, ...), that means either a Planning agent needs to then realize that and recursively generate more vector search queries (mini-AutoGPT)... or we need to index on more than the sentence embedding.
Again, super interesting problems, and we're hiring for folks interested in helping work on it!